-
Notifications
You must be signed in to change notification settings - Fork 2.9k
Pull requests: JustVugg/colibri
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
qwen36: run Swiftlet qpack containers end to end (dense from safetensors, routed experts via bounded Metal slots)
#1323
opened Sep 2, 2026 by
Avicennasis
Loading…
glm53: size the expert cache around the model, not around MemAvailable
#1321
opened Sep 2, 2026 by
ronaldcklomp
Loading…
fix(build): fall back to the Homebrew prefixes when brew is off the PATH
#1320
opened Sep 2, 2026 by
Avicennasis
Loading…
fix(qwen36): read tokenizer.json merges in string and pair spellings
#1319
opened Sep 2, 2026 by
Avicennasis
Loading…
fix(planner): honor cgroup memory limits
#1316
opened Sep 1, 2026 by
Avicennasis
Loading…
3 of 5 tasks
Add resumable, hash-verified qpack installers for Hugging Face and static mirrors
#1315
opened Sep 1, 2026 by
Avicennasis
Loading…
perf(quant): compute four output rows per pass in matmul_fp8 (+21.34% tok/s, bit-exact)
#1313
opened Sep 1, 2026 by
BrianHeeseIs
Loading…
fix(coli): discover serve processes where there is no /proc
#1307
opened Aug 31, 2026 by
BrianHeeseIs
Loading…
docs: add a reproducible benchmarking protocol
#1294
opened Aug 31, 2026 by
ZacharyZcR
Contributor
Loading…
docs: lead Windows DeepSeek users to release launcher
#1291
opened Aug 31, 2026 by
ZacharyZcR
Contributor
Loading…
Add MLX affine qpack reader, inspector, and packed affine Metal GEMV
#1290
opened Aug 31, 2026 by
Avicennasis
Loading…
build: list the headers each engine includes as Makefile prerequisites
#1284
opened Aug 30, 2026 by
Avicennasis
Loading…
feat(xdna): optional explicit AMD XDNA2 lane for the GLM shared expert
#1261
opened Aug 28, 2026 by
Kenneth-Javier
Contributor
Loading…
fix(cuda): charge the expert tier real VRAM, not logical bytes (#687)
#1244
opened Aug 26, 2026 by
Unknown-Findout
Contributor
Loading…
perf(sse41): add SSE4.1 fallback tier for olmoe and qwen36 int8 GEMV
#1239
opened Aug 26, 2026 by
jtinbergen
Loading…
feat(experiments): validate reproducible performance records
#1238
opened Aug 26, 2026 by
ZacharyZcR
Contributor
Loading…
feat(glm52): optimize inference on single RTX 3090 (24GB) via Direct I/O and MTP
#1177
opened Aug 22, 2026 by
up1t3
Loading…
feat: fmt=8 (fp8-e4m3) decode on the kv_b absorb path, CPU and CUDA
#1102
opened Aug 19, 2026 by
monotophic
Contributor
Loading…
cuda: zero-copy expert views on pageable-shared memory (GB10)
#936
opened Aug 11, 2026 by
Nanetnounou
Contributor
Loading…
cuda: warp-per-row E8 kernels, lattice table in shared memory
#935
opened Aug 11, 2026 by
Nanetnounou
Contributor
Loading…
Previous Next
ProTip!
Find all pull requests that aren't related to any open issues with -linked:issue.