Back to Tools Directory

PaperGuru-Benchmark

TeX

Lifecycle-Aware Memory for long-horizon LLM agents — 66.05% on PaperBench, 94.66% on SurveyBench, 10 peer-reviewed acceptances at FSE/ICML/TOSEM/AEI/ICoGB

Categories & Topics

llm-agentsbenchmarklong-horizonmemory
View on GitHub

Repository Stats

1.3k
GitHub Stars
197
Forks
1.3k
Watchers
0
Open Issues
License
Other
Repository
PaperGuru-AI/PaperGuru-Benchmark
Last Updated
Today
August 7, 2026
Top Contributors
PaperGuru
9
paperguruai
4
View all contributors →

Related Tools

agentic-stack

2.2k

One brain, many harnesses. Portable .agent/ folder (memory + skills + protocols) that plugs into Claude Code, Cursor, Windsurf, OpenCode, OpenClaw, Hermes, or DIY Python — and keeps its knowledge when you switch.

AlayaWorld

841

Full-stack open-source interactive long-horizon world model.

awesome-ai-agent-papers

1.7k

A curated collection of AI agent research papers released in 2026, covering agent engineering, memory, evaluation, workflows, and autonomous systems.

boop-agent

1.3k

iMessage personal agent: choose Claude Agent SDK (Claude Code) or Codex app-server runtime (Codex/ChatGPT), with memory, sub-agents, automations, integrations.

deep-swe

1.3k

Measuring frontier coding agents on original, long-horizon engineering tasks

Engram

4.6k

Engram enhances memory capabilities in large language models by using a scalable lookup method. This approach improves performance and efficiency, making it easier to access and utilize conditional information for various applications.