Skip to content

Pull requests: JustVugg/colibri

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

metal: GPU attention for S>4 prefill (single command buffer)
#763 opened Aug 1, 2026 by RDouglasSharp Contributor Loading…
inkling: shared experts to the GPU, 1.75 to 2.46 tok/s on Apple Silicon metal Backend Metal/Apple performance Velocità / tok-s / ottimizzazioni
#757 opened Aug 1, 2026 by rgbkrk Contributor Loading…
feat(win): fix silent CPU fallback, launcher suite, DirectStorage expert loads enhancement New feature or request needs-rebase Confligge, serve rebase dell'autore
#670 opened Jul 28, 2026 by khalilswdp Contributor Loading…
8 tasks done
feat(core): model-architecture seam, chat templates, text antiprompt enhancement New feature or request needs-rebase Confligge, serve rebase dell'autore
#667 opened Jul 28, 2026 by khalilswdp Contributor Loading…
5 tasks done
QLoRA training path: fine-tune GLM-5.2 (744B) in 64 GB RAM enhancement New feature or request feature Nuova funzionalità
#626 opened Jul 26, 2026 by pavolbauer Loading…
5 tasks done
Metal fmt=4 grouped-int4 decode: attention + routed experts (#585) metal Backend Metal/Apple needs-rebase Confligge, serve rebase dell'autore
#587 opened Jul 24, 2026 by RDouglasSharp Contributor Loading…
Preserve model-declared EOS tokens in serve mode enhancement New feature or request needs-rebase Confligge, serve rebase dell'autore
#584 opened Jul 24, 2026 by saskw2010 Draft
Fix stateful KV tail at the NGEN limit bug Difetto verificato nel codice
#567 opened Jul 23, 2026 by winklemad Contributor Loading…
3 of 5 tasks
add persistence controls for lower SSD writes enhancement New feature or request
#555 opened Jul 23, 2026 by Skater1808 Loading…
5 tasks
CPU: KV cache quantization — KV8 (fp8 e4m3) + KV_TQ (rotated-int4 / PolarQuant) enhancement New feature or request performance Velocità / tok-s / ottimizzazioni
#553 opened Jul 23, 2026 by NeuralNotwerk Contributor Loading…
feat: add distributed expert workers enhancement New feature or request
#551 opened Jul 23, 2026 by gauravsaini Loading…
feat: add dense MLP activation sharding enhancement New feature or request performance Velocità / tok-s / ottimizzazioni
#550 opened Jul 23, 2026 by gauravsaini Loading…
feat: add Qwen3-30B-A3B engine model-support Supporto a nuovi modelli needs-rebase Confligge, serve rebase dell'autore
#544 opened Jul 23, 2026 by opxyc Contributor Loading…
4 of 5 tasks
pilot: multi-worker PILOT_REAL prefetch (PILOT_WORKERS) — byte-identical base for the #441 hardware A/B performance Velocità / tok-s / ottimizzazioni
#480 opened Jul 21, 2026 by cdhdt Contributor Draft
4 of 5 tasks
nix: CUDA support enhancement New feature or request
#416 opened Jul 19, 2026 by attilaolah Contributor Draft
5 tasks
KV cache quantization: fp8 (KV8) + 4-bit TurboQuant (KV_TQ) on CPU, CUDA, and Metal enhancement New feature or request performance Velocità / tok-s / ottimizzazioni
#399 opened Jul 18, 2026 by NeuralNotwerk Contributor Loading…
feat: Add NUMA-aware RAM-disk streaming discussion Proposta / discussione aperta, non un task enhancement New feature or request needs-rebase Confligge, serve rebase dell'autore performance Velocità / tok-s / ottimizzazioni
#377 opened Jul 17, 2026 by BColsey Loading…
Nearly double the hit-rate needs-rebase Confligge, serve rebase dell'autore performance Velocità / tok-s / ottimizzazioni
#223 opened Jul 14, 2026 by withinboredom Contributor Loading…
4 of 5 tasks
ProTip! Type g i on any issue or pull request to go back to the issue listing page.