-
Notifications
You must be signed in to change notification settings - Fork 2.3k
Pull requests: JustVugg/colibri
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
inkling: shared experts to the GPU, 1.75 to 2.46 tok/s on Apple Silicon
#757
opened Aug 1, 2026 by
rgbkrk
Contributor
Loading…
docs: inkling audio input and the 10 MB sidecar fetch
#756
opened Aug 1, 2026 by
rgbkrk
Contributor
Loading…
build: Linux/aarch64 branch in the Makefile — probe DOTPROD/I8MM instead of trusting native (#631)
#755
opened Aug 1, 2026 by
anrasi
Loading…
olmoe: adopt route_trace.h — the last engine on the shared history
#754
opened Aug 1, 2026 by
terrizoaguimor
Contributor
Loading…
Vulkan MoE GEMV backend for integrated/AMD GPUs (draft, complements #418)
enhancement
New feature or request
vulkan
Backend Vulkan/AMD
feat(qwen36): CUDA VRAM expert tier — heat-based placement across GPUs via the shared CUDA backend
cuda
Backend CUDA/NVIDIA
model-support
Supporto a nuovi modelli
feat(qwen36): Qwen3.6-35B-A3B engine (CPU): hybrid Gated Attention + Gated DeltaNet + streaming MoE
model-support
Supporto a nuovi modelli
#712
opened Jul 30, 2026 by
kreuzzelg
Loading…
feat(win): fix silent CPU fallback, launcher suite, DirectStorage expert loads
enhancement
New feature or request
needs-rebase
Confligge, serve rebase dell'autore
#670
opened Jul 28, 2026 by
khalilswdp
Contributor
Loading…
8 tasks done
feat(core): model-architecture seam, chat templates, text antiprompt
enhancement
New feature or request
needs-rebase
Confligge, serve rebase dell'autore
#667
opened Jul 28, 2026 by
khalilswdp
Contributor
Loading…
5 tasks done
QLoRA training path: fine-tune GLM-5.2 (744B) in 64 GB RAM
enhancement
New feature or request
feature
Nuova funzionalità
#626
opened Jul 26, 2026 by
pavolbauer
Loading…
5 tasks done
WIP: MiniMax-M3 support — GQA + MSA block-sparse attention, o200k tokenizer, converter (follow-up to #418)
model-support
Supporto a nuovi modelli
#601
opened Jul 24, 2026 by
steve-m
Contributor
Loading…
5 tasks
Metal fmt=4 grouped-int4 decode: attention + routed experts (#585)
metal
Backend Metal/Apple
needs-rebase
Confligge, serve rebase dell'autore
#587
opened Jul 24, 2026 by
RDouglasSharp
Contributor
Loading…
Preserve model-declared EOS tokens in serve mode
enhancement
New feature or request
needs-rebase
Confligge, serve rebase dell'autore
Fix stateful KV tail at the NGEN limit
bug
Difetto verificato nel codice
#567
opened Jul 23, 2026 by
winklemad
Contributor
Loading…
3 of 5 tasks
add persistence controls for lower SSD writes
enhancement
New feature or request
#555
opened Jul 23, 2026 by
Skater1808
Loading…
5 tasks
CPU: KV cache quantization — KV8 (fp8 e4m3) + KV_TQ (rotated-int4 / PolarQuant)
enhancement
New feature or request
performance
Velocità / tok-s / ottimizzazioni
#553
opened Jul 23, 2026 by
NeuralNotwerk
Contributor
Loading…
feat: add distributed expert workers
enhancement
New feature or request
#551
opened Jul 23, 2026 by
gauravsaini
Loading…
feat: add dense MLP activation sharding
enhancement
New feature or request
performance
Velocità / tok-s / ottimizzazioni
#550
opened Jul 23, 2026 by
gauravsaini
Loading…
feat: add Qwen3-30B-A3B engine
model-support
Supporto a nuovi modelli
needs-rebase
Confligge, serve rebase dell'autore
#544
opened Jul 23, 2026 by
opxyc
Contributor
Loading…
4 of 5 tasks
pilot: multi-worker PILOT_REAL prefetch (PILOT_WORKERS) — byte-identical base for the #441 hardware A/B
performance
Velocità / tok-s / ottimizzazioni
Port the coli CLI launcher to Go (dependency-free) — first step for #310
enhancement
New feature or request
nix: CUDA support
enhancement
New feature or request
#416
opened Jul 19, 2026 by
attilaolah
Contributor
•
Draft
5 tasks
KV cache quantization: fp8 (KV8) + 4-bit TurboQuant (KV_TQ) on CPU, CUDA, and Metal
enhancement
New feature or request
performance
Velocità / tok-s / ottimizzazioni
#399
opened Jul 18, 2026 by
NeuralNotwerk
Contributor
Loading…
feat: Add NUMA-aware RAM-disk streaming
discussion
Proposta / discussione aperta, non un task
enhancement
New feature or request
needs-rebase
Confligge, serve rebase dell'autore
performance
Velocità / tok-s / ottimizzazioni
#377
opened Jul 17, 2026 by
BColsey
Loading…
Nearly double the hit-rate
needs-rebase
Confligge, serve rebase dell'autore
performance
Velocità / tok-s / ottimizzazioni
#223
opened Jul 14, 2026 by
withinboredom
Contributor
Loading…
4 of 5 tasks
Previous Next
ProTip!
Exclude everything labeled
bug with -label:bug.