A privacy-first, beginner-friendly desktop application for running GGUF language models fully locally on Windows, Linux, and macOS. No data ever leaves your machine.
โ ๏ธ This is a testing/preview release. It is locked to one model while we validate the new plan-driven agent architecture. Expect rough edges. Report issues you find.
v2.3.0 is a testing release that takes the agentic experience further with a strictly plan-driven agent and a developer-style inline process UI โ but it is temporarily pinned to a single model (Gemma 4 12B Instruct Q4_K_M) to tune the agent's performance before expanding model support.
- ๐ฏ Single-model focus (temporary) โ optimized for Gemma 4 12B Instruct Q4_K_M; multi-model detection removed so this one model just works. Universal loading returns next release.
- ๐ Strictly plan-driven agent โ every turn goes through a planner node: tool-free questions answered directly; tasks get a step-by-step plan executed with sandboxed tools. The reactive ReAct fallback was removed.
- ๐ฅ๏ธ Developer-style process UI โ plan steps, tool calls, and results render inline in the chat above each answer.
- ๐งน Reliable final answers โ stray tool-call JSON and stale status text are scrubbed from replies.
- ๐ฅ Auto-load at startup โ scans
models/folder and loads the pinned GGUF automatically; missing model downloads with live progress. - ๐ง Pruned tools โ 10 focused workspace tools (memory/meta tools removed).
- ๐ Self-setup launchers โ
launch.bat/launch.shhandle all dependency installation.
| Artifact | Size | Notes |
|---|---|---|
| GGUFLoader_v2.3.0_CPU.exe | ~145 MB | Windows ยท CPU-only, works everywhere |
| GGUFLoader_v2.3.0_GPU.exe | ~930 MB | Windows ยท NVIDIA CUDA |
| GGUFLoader_v2.3.0_linux_x86_64_CPU | ~50 MB | Linux ยท CPU-only |
Click any filename above to download directly. The CUDA build bundles the full CUDA runtime; the CPU build is several times smaller.
- Start the app โ it auto-loads the pinned Gemma 4 12B Q4_K_M from the
models/folder in the background. - No model on disk? The model chip in the header downloads it with live progress and loads it when finished.
- Chat in the main window, or press Ctrl/Cmd + Shift + A for Agent Mode and choose a workspace folder.
Universal Model Loader returns โ the next release removes the single-model restriction. You'll be able to run any GGUF model with the agent, with hardware-aware recommendations so you can pick a model that fits your PC's RAM/VRAM.
Install via pip:
pip install ggufloaderthen runggufloaderView on PyPI
GGUF Loader v2.2.0 is the stable, universal model loader. You pick any GGUF model that fits your PC's resources โ from small 3B models on modest machines to large 70B+ models on high-end rigs. The app detects your hardware (RAM, VRAM, CPU/GPU) and recommends models that work for you.
- ๐ Universal model loader โ load any GGUF model, no restrictions
- ๐ฆ PyPI package โ
pip install ggufloader, works in any Python environment - ๐ค Basic agentic mode โ LangGraph-driven agent with 7 sandboxed tools
- ๐ Find Paragraph โ locate a passage in a document or folder, no RAG needed
- ๐ Full-folder summaries โ reads every readable file before answering
- โก One-click GPU โ install CUDA/Metal support from the Settings UI
- ๐ฌ Streaming chat โ token-by-token delivery via WebSocket
- ๐ 100% local โ no cloud, no subscriptions, no data leaves your machine
Option 1: pip
pip install ggufloader
ggufloaderOption 2: Prebuilt executable
| Artifact | Size | Notes |
|---|---|---|
| GGUFLoader_v2.2.0_GPU.exe | ~850 MB | Windows ยท NVIDIA CUDA |
| GGUFLoader_v2.2.0_CPU.exe | ~70 MB | Windows ยท CPU-only |
| GGUFLoader_v2.2.0_linux_x86_64_CPU | ~105 MB | Linux ยท CPU-only |
Click any filename above to download directly.
- ๐๏ธ Sessions left, chat center, tools right โ sessions (left), chat with inline agent process (center), tool panels (right)
- ๐ฌ Streaming chat โ token-by-token delivery via WebSocket
- ๐ค Agent mode โ planner-driven runs with Allow/Deny approvals, streamed inline in the chat
- ๐ฅ Model chip โ shows load state; downloads the model with live progress
- โ๏ธ Settings โ Model, Providers, Agent, Hardware, Appearance, Keyboard, Plugins tabs
- ๐จ Theme โ dark/light mode, accent colors, font size
- โจ๏ธ Command palette โ Ctrl/Cmd+K, Ctrl/Cmd+Shift+A (agent mode)
- ๐ฑ Electron โ standalone desktop app (no browser needed)
- ๐ค Plan-driven agent (LangGraph) โ a planner node decides each turn: tool-free questions answered directly; tasks get a step-by-step plan with sandboxed tools inside your workspace, with Allow/Deny approval for commands, code, and git writes, and SQLite checkpointing.
- ๐ Advanced Search (Find Paragraph) โ locate a passage in a document or folder with the model itself, no RAG or vector database required.
- ๐งพ Real file reading โ extracts text from
.md,.pdf,.docx,.txtand source files. - โก GPU acceleration โ enable under Settings โ Hardware (installs CUDA/Metal build with live status).
- ๐ Privacy first โ 100% local inference. Your prompts and files never leave your machine.
- ๐ป Cross-platform โ Windows 10/11, Linux, and macOS (including Apple Silicon), via React + Electron or in the browser.
Agentic Mode turns the local model into a working assistant for a folder you choose. It plans multi-step tasks, calls tools, and streams every step live.
| Tool | What it does |
|---|---|
list_directory / glob |
Explore folders and match file paths |
read_file |
Read any file (MD/PDF/DOCX/TXT/code) |
search_files |
Find files and grep for content |
write_file / edit_file / move_file |
Create, edit, and move files |
run_command |
Run a shell command (approval-gated) |
run_python |
Execute Python source (approval-gated) |
git |
Git operations (writes require approval) |
Shell commands, code execution, and git writes pause for an Allow / Deny card. Everything else (reading, searching, file edits) runs automatically.
- "Summarize the Day 4 folder" โ reads all files and gives a real summary
- "Create a new feature module with proper structure"
- "Refactor this codebase and organize files"
- "Find where
MAX_TOKENSis defined and explain it"
Open the Advanced Search panel to locate a specific passage:
- Single file โ type a question and the model finds matching paragraphs
- Folder search โ a planner decides which files to read, with live progress
- Smart defaults โ your last query, source, and settings are remembered
No vector database, no embeddings โ just the model reading the text.
The app runs on CPU by default. To enable GPU:
- Open Settings โ Hardware and click the GPU install option.
- The app installs the CUDA-enabled build and shows a green tick when done.
- Restart the app โ inference now uses the GPU.
Manual scripts: scripts/install_gpu_llama.bat (Windows) /
scripts/install_gpu_llama.sh (Linux/macOS).
- OS: Windows 10/11, Linux, macOS (Intel & Apple Silicon)
- RAM: 32 GB minimum recommended
- Storage: ~8 GB free for the model file
- GPU: Optional โ NVIDIA CUDA on Windows/Linux, Metal on macOS
- Python: 3.10โ3.13 (for running from source only)
- Windows executable:
scripts/build_exe.batโ detects CUDA and names the outputGGUFLoader_v<version>_GPU.exeor_CPU.exeautomatically. - Linux executable:
scripts/build_linux.shโ must run on Linux (or WSL); producesGGUFLoader_v<version>_linux_x86_64_CPU. - Tests:
pip install pytest && python -m pytest
- Quick Reference
- Developing Addons
- AGENTS.md โ codebase guide for AI coding agents
- Architecture
- Changelog
- Contributing
- Security Policy
Contributions are welcome! See CONTRIBUTING.md.
MIT โ see LICENSE.
Report vulnerabilities to hussainnazary475@gmail.com or see SECURITY.md.
- ๐ Report Issues
- ๐ฌ Discussions
- ๐ง hussainnazary475@gmail.com
Built with โค๏ธ by the GGUF Loader community


