Skip to content
ย 
ย 

Latest commit

ย 

History

244 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

GGUF Loader

GitHub License GitHub Last Commit

A privacy-first, beginner-friendly desktop application for running GGUF language models fully locally on Windows, Linux, and macOS. No data ever leaves your machine.


๐Ÿงช v2.3.0 โ€” Testing Release (Single-Model Agent)

โš ๏ธ This is a testing/preview release. It is locked to one model while we validate the new plan-driven agent architecture. Expect rough edges. Report issues you find.

v2.3.0 is a testing release that takes the agentic experience further with a strictly plan-driven agent and a developer-style inline process UI โ€” but it is temporarily pinned to a single model (Gemma 4 12B Instruct Q4_K_M) to tune the agent's performance before expanding model support.

What's new in v2.3.0

  • ๐ŸŽฏ Single-model focus (temporary) โ€” optimized for Gemma 4 12B Instruct Q4_K_M; multi-model detection removed so this one model just works. Universal loading returns next release.
  • ๐Ÿ“‹ Strictly plan-driven agent โ€” every turn goes through a planner node: tool-free questions answered directly; tasks get a step-by-step plan executed with sandboxed tools. The reactive ReAct fallback was removed.
  • ๐Ÿ–ฅ๏ธ Developer-style process UI โ€” plan steps, tool calls, and results render inline in the chat above each answer.
  • ๐Ÿงน Reliable final answers โ€” stray tool-call JSON and stale status text are scrubbed from replies.
  • ๐Ÿ“ฅ Auto-load at startup โ€” scans models/ folder and loads the pinned GGUF automatically; missing model downloads with live progress.
  • ๐Ÿ”ง Pruned tools โ€” 10 focused workspace tools (memory/meta tools removed).
  • ๐Ÿš€ Self-setup launchers โ€” launch.bat / launch.sh handle all dependency installation.

Screenshots (v2.3.0)

v2.3.0 - Chat Interface

v2.3.0 - Agent Mode

v2.3.0 - Settings

Download v2.3.0

Artifact Size Notes
GGUFLoader_v2.3.0_CPU.exe ~145 MB Windows ยท CPU-only, works everywhere
GGUFLoader_v2.3.0_GPU.exe ~930 MB Windows ยท NVIDIA CUDA
GGUFLoader_v2.3.0_linux_x86_64_CPU ~50 MB Linux ยท CPU-only

Click any filename above to download directly. The CUDA build bundles the full CUDA runtime; the CPU build is several times smaller.

First launch (v2.3.0)

  1. Start the app โ€” it auto-loads the pinned Gemma 4 12B Q4_K_M from the models/ folder in the background.
  2. No model on disk? The model chip in the header downloads it with live progress and loads it when finished.
  3. Chat in the main window, or press Ctrl/Cmd + Shift + A for Agent Mode and choose a workspace folder.

๐Ÿ”ฎ What's Coming Next

Universal Model Loader returns โ€” the next release removes the single-model restriction. You'll be able to run any GGUF model with the agent, with hardware-aware recommendations so you can pick a model that fits your PC's RAM/VRAM.


๐Ÿ“ฆ v2.2.0 โ€” Stable Release (Universal Model Loader)

Install via pip: pip install ggufloader then run ggufloader View on PyPI

GGUF Loader v2.2.0 is the stable, universal model loader. You pick any GGUF model that fits your PC's resources โ€” from small 3B models on modest machines to large 70B+ models on high-end rigs. The app detects your hardware (RAM, VRAM, CPU/GPU) and recommends models that work for you.

What v2.2.0 includes

  • ๐ŸŒ Universal model loader โ€” load any GGUF model, no restrictions
  • ๐Ÿ“ฆ PyPI package โ€” pip install ggufloader, works in any Python environment
  • ๐Ÿค– Basic agentic mode โ€” LangGraph-driven agent with 7 sandboxed tools
  • ๐Ÿ”Ž Find Paragraph โ€” locate a passage in a document or folder, no RAG needed
  • ๐Ÿ“‚ Full-folder summaries โ€” reads every readable file before answering
  • โšก One-click GPU โ€” install CUDA/Metal support from the Settings UI
  • ๐Ÿ’ฌ Streaming chat โ€” token-by-token delivery via WebSocket
  • ๐Ÿ”’ 100% local โ€” no cloud, no subscriptions, no data leaves your machine

Install v2.2.0

Option 1: pip

pip install ggufloader
ggufloader

Option 2: Prebuilt executable

Artifact Size Notes
GGUFLoader_v2.2.0_GPU.exe ~850 MB Windows ยท NVIDIA CUDA
GGUFLoader_v2.2.0_CPU.exe ~70 MB Windows ยท CPU-only
GGUFLoader_v2.2.0_linux_x86_64_CPU ~105 MB Linux ยท CPU-only

Click any filename above to download directly.


โœจ Features

Interface

  • ๐Ÿ—‚๏ธ Sessions left, chat center, tools right โ€” sessions (left), chat with inline agent process (center), tool panels (right)
  • ๐Ÿ’ฌ Streaming chat โ€” token-by-token delivery via WebSocket
  • ๐Ÿค– Agent mode โ€” planner-driven runs with Allow/Deny approvals, streamed inline in the chat
  • ๐Ÿ“ฅ Model chip โ€” shows load state; downloads the model with live progress
  • โš™๏ธ Settings โ€” Model, Providers, Agent, Hardware, Appearance, Keyboard, Plugins tabs
  • ๐ŸŽจ Theme โ€” dark/light mode, accent colors, font size
  • โŒจ๏ธ Command palette โ€” Ctrl/Cmd+K, Ctrl/Cmd+Shift+A (agent mode)
  • ๐Ÿ“ฑ Electron โ€” standalone desktop app (no browser needed)

Core Features

  • ๐Ÿค– Plan-driven agent (LangGraph) โ€” a planner node decides each turn: tool-free questions answered directly; tasks get a step-by-step plan with sandboxed tools inside your workspace, with Allow/Deny approval for commands, code, and git writes, and SQLite checkpointing.
  • ๐Ÿ”Ž Advanced Search (Find Paragraph) โ€” locate a passage in a document or folder with the model itself, no RAG or vector database required.
  • ๐Ÿงพ Real file reading โ€” extracts text from .md, .pdf, .docx, .txt and source files.
  • โšก GPU acceleration โ€” enable under Settings โ†’ Hardware (installs CUDA/Metal build with live status).
  • ๐Ÿ”’ Privacy first โ€” 100% local inference. Your prompts and files never leave your machine.
  • ๐Ÿ’ป Cross-platform โ€” Windows 10/11, Linux, and macOS (including Apple Silicon), via React + Electron or in the browser.

๐Ÿค– Agentic Mode

Agentic Mode turns the local model into a working assistant for a folder you choose. It plans multi-step tasks, calls tools, and streams every step live.

Tools

Tool What it does
list_directory / glob Explore folders and match file paths
read_file Read any file (MD/PDF/DOCX/TXT/code)
search_files Find files and grep for content
write_file / edit_file / move_file Create, edit, and move files
run_command Run a shell command (approval-gated)
run_python Execute Python source (approval-gated)
git Git operations (writes require approval)

Human approval

Shell commands, code execution, and git writes pause for an Allow / Deny card. Everything else (reading, searching, file edits) runs automatically.

Example tasks

  • "Summarize the Day 4 folder" โ†’ reads all files and gives a real summary
  • "Create a new feature module with proper structure"
  • "Refactor this codebase and organize files"
  • "Find where MAX_TOKENS is defined and explain it"

๐Ÿ”Ž Advanced Search (Find Paragraph, no RAG)

Open the Advanced Search panel to locate a specific passage:

  • Single file โ€” type a question and the model finds matching paragraphs
  • Folder search โ€” a planner decides which files to read, with live progress
  • Smart defaults โ€” your last query, source, and settings are remembered

No vector database, no embeddings โ€” just the model reading the text.


โšก GPU Acceleration

The app runs on CPU by default. To enable GPU:

  1. Open Settings โ†’ Hardware and click the GPU install option.
  2. The app installs the CUDA-enabled build and shows a green tick when done.
  3. Restart the app โ€” inference now uses the GPU.

Manual scripts: scripts/install_gpu_llama.bat (Windows) / scripts/install_gpu_llama.sh (Linux/macOS).


๐Ÿ› ๏ธ System Requirements

  • OS: Windows 10/11, Linux, macOS (Intel & Apple Silicon)
  • RAM: 32 GB minimum recommended
  • Storage: ~8 GB free for the model file
  • GPU: Optional โ€” NVIDIA CUDA on Windows/Linux, Metal on macOS
  • Python: 3.10โ€“3.13 (for running from source only)

๐Ÿงฑ Building from source

  • Windows executable: scripts/build_exe.bat โ€” detects CUDA and names the output GGUFLoader_v<version>_GPU.exe or _CPU.exe automatically.
  • Linux executable: scripts/build_linux.sh โ€” must run on Linux (or WSL); produces GGUFLoader_v<version>_linux_x86_64_CPU.
  • Tests: pip install pytest && python -m pytest

๐Ÿ“š Documentation


๐Ÿค Contributing

Contributions are welcome! See CONTRIBUTING.md.

๐Ÿ“„ License

MIT โ€” see LICENSE.

๐Ÿ”’ Security

Report vulnerabilities to hussainnazary475@gmail.com or see SECURITY.md.

๐Ÿ“ž Support


Built with โค๏ธ by the GGUF Loader community

About

GGUF Loader with its Agentic Mode, and floating button, ai Models | Open Source & Offline. Mistral, Deepseek, llama, gemma, qwen

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages