Back to Tools Directory

turboquant-pytorch

Python

From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% attention fidelity.

Categories & Topics

llmkv-cache-compressionpytorchquantization
View on GitHub

Repository Stats

1.0k
GitHub Stars
138
Forks
1.0k
Watchers
17
Open Issues
License
MIT License
Repository
tonbistudio/turboquant-pytorch
Last Updated
Yesterday
August 6, 2026
Top Contributors
tonbistudio
9
View all contributors →

Related Tools

abliterix

189

Automated alignment adjustment for LLMs — direct steering, LoRA, and MoE expert-granular abliteration, optimized via multi-objective Optuna TPE.

agent-os

4.3k

Give agents an operating system as a library. Runs in your existing backend – no sandboxes, VMs, or SaaS. Powered by WebAssembly & V8 isolates.

ainovel-cli

1.6k

✨多agent实现全自动AI小说生成

astrid

10k

Astrid is a portable, capability-secure operating system for composable software.

awesome-ai-agent-papers

1.7k

A curated collection of AI agent research papers released in 2026, covering agent engineering, memory, evaluation, workflows, and autonomous systems.

awesome-ai-apps

13k

A collection of projects showcasing RAG, agents, workflows, and other AI use cases