Skip to content
#

quantisation

Here are 22 public repositories matching this topic...

LoRA fine-tune and serve NVFP4 models on one DGX Spark (GB10, 128 GB UMA): text backbones via generic-family onboarding, plus VLMs (vision tower, or LLM+tower jointly via --train-target both) validated end-to-end on Pixtral and Nemotron-Omni. Fused Triton dequant; runtime-LoRA and merge serving.

  • Updated Aug 1, 2026
  • Python

Do compressed classroom LLMs keep their answers, not just their scores? Fidelity, confident error, semantic entropy and measured inference energy across three model families at three quantisation levels.

  • Updated Sep 30, 2026
  • Python

Add this topic to your repo

To associate your repository with the quantisation topic, visit your repo's landing page and select "manage topics."

Learn more