building world's fastest retrieval & inference algorithms
you can visit www dot srswti dot com to know more.
building world's fastest retrieval & inference algorithms
you can visit www dot srswti dot com to know more.
Native Intel Xe2 LLM inference: SYCL/DPAS kernels, paged KV, prefix caching, continuous batching, and XMX packed prefill
bare-cuda fp6 inference engine for qwen3.8-27b on one rtx 5090 — sm120 tensor-core kernels, 256k context, near-lossless perf too!
nvfp4 gemm specialized for llm decode on consumer blackwell (sm120). 1.42x over cutlass 4.8 at batch 1, bit-exact.
run a 118b moe on one 5090 and two intel b70s. 88 tok/s. yes cuda and sycl both runtimes talking to each other. beating rtx pro 6000 on 120k ctx plus.
A high-throughput and memory-efficient inference and serving engine for LLMs
gpu-native kv cache compression and streaming infrastructure for long-context inference. ummm basically, nccl but through ethernet
axe - a precision agentic coder. large codebases. zero bloat. terminal-native. precise retrieval. powerful inference.
This organization has no public members. You must be a member to see who’s a part of this organization.
Loading…
Loading…