A powerful model compression framework for LLMs and LVLMs, adapted for NVIDIA GPUs and Huawei Ascend NPUs.
Categories & Topics
Repository Stats
Related Tools
apexyard
Your AI engineering org, governed. Ship agent-built code safely to production.
auto-harness
Bring your own agent and build a self-improving agentic system. Automatically mine failures, optimize the agent harness, and gate against regressions.
bullshit-benchmark
BullshitBench evaluates how AI models respond to nonsensical prompts, determining whether they challenge the prompts instead of providing confident answers. This helps improve the reliability and understanding of AI responses in ambiguous situations.
darwin-skill
达尔文.skill —— 一个让你的Skill无限进化的系统:评估→改进→测试→保留或回滚 | Autoresearch-inspired autonomous skill optimization for Claude Code. Evaluate, improve, test, keep or revert.
deep-swe
Measuring frontier coding agents on original, long-horizon engineering tasks
opik
Monitor and improve the performance of large language model applications with easy-to-use dashboards that provide insights, automated evaluations, and detailed tracing. This tool helps users debug and assess their systems to ensure optimal functionality and reliability.