From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% attention fidelity.
Categories & Topics
Repository Stats
Related Tools
abliterix
Automated alignment adjustment for LLMs — direct steering, LoRA, and MoE expert-granular abliteration, optimized via multi-objective Optuna TPE.
agent-os
Give agents an operating system as a library. Runs in your existing backend – no sandboxes, VMs, or SaaS. Powered by WebAssembly & V8 isolates.
ainovel-cli
✨多agent实现全自动AI小说生成
astrid
Astrid is a portable, capability-secure operating system for composable software.
awesome-ai-agent-papers
A curated collection of AI agent research papers released in 2026, covering agent engineering, memory, evaluation, workflows, and autonomous systems.
awesome-ai-apps
A collection of projects showcasing RAG, agents, workflows, and other AI use cases