Skip to content
FrameworkHunt

Agent frameworks, evaluated against evidence you can check.

Every score here points at a file, at a pinned commit, containing the thing it claims. The build fails if the citation does not hold. Popularity is reported separately and never touches the score.

4

frameworks evaluated

79

verified citations

17

research records

Rubric v1.0.0 · citations verified 18 Aug 2026 · repository snapshot 18 Aug 2026

Ranked by architecture score

How scoring works →
95/100

Pydantic AI

Pydantic · Python · pydantic/pydantic-ai

Type-first agent framework that treats durability and OpenTelemetry as architecture, not integrations.

  • control 3
  • state 4
  • tools 4
  • reliability 4
  • observability 4
  • developer 4
89/100

LangGraph

LangChain · Python · langchain-ai/langgraph

Low-level graph runtime for long-running, stateful agents, with checkpointing as a first-class primitive.

  • control 4
  • state 4
  • tools 3
  • reliability 4
  • observability 3
  • developer 3
86/100

OpenAI Agents SDK

OpenAI · Python · openai/openai-agents-python

Small primitive set — agents, handoffs, guardrails, sessions — with tracing built in rather than bolted on.

  • control 3
  • state 3
  • tools 4
  • reliability 4
  • observability 4
  • developer 3
80/100

CrewAI

CrewAI Inc. · Python · crewAIInc/crewAI

Role-based crews for delegation-shaped work, plus a decorator flow layer and event-driven checkpointing.

  • control 3
  • state 4
  • tools 3
  • reliability 3
  • observability 3
  • developer 3

Two layers, never blended

An architecture score built only from evidenced axes, and a repository snapshot shown beside it. Stars cannot make a design better. The score function is structurally unable to read them.

Citations pinned to a commit

Each claim names a path and the strings that must appear in it. A verifier fetches that file at the reviewed SHA and fails the build when the text is missing.

Unknown is not zero

An axis without sufficient evidence stays Unknown, and a record with any Unknown axis publishes no aggregate at all rather than a flattering guess.

Papers, memos and threads kept deliberately separate from the directory. They are not frameworks, carry no score, and are never counted as evaluated tools.