This repository contains the code to train and evaluate TRIBE v2, a multimodal model for brain response prediction
Categories & Topics
Repository Stats
Related Tools
auto-harness
Bring your own agent and build a self-improving agentic system. Automatically mine failures, optimize the agent harness, and gate against regressions.
bullshit-benchmark
BullshitBench evaluates how AI models respond to nonsensical prompts, determining whether they challenge the prompts instead of providing confident answers. This helps improve the reliability and understanding of AI responses in ambiguous situations.
darwin-skill
达尔文.skill —— 一个让你的Skill无限进化的系统:评估→改进→测试→保留或回滚 | Autoresearch-inspired autonomous skill optimization for Claude Code. Evaluate, improve, test, keep or revert.
deep-swe
Measuring frontier coding agents on original, long-horizon engineering tasks
DeepSpec
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
InterviewOS
Replace coding puzzles with real-work simulations.