Long-Horizon Agents
Research track · evidence through September 24, 2026

When does an agent become long-horizon?

Not when it produces a long transcript. A long-horizon agent must preserve goals and state, coordinate dependent actions, detect drift, recover from failure, and keep its evidence usable after the original context window is gone.

Living research track Research cutoff Primary papers and official project reports.
49primary & official sources
15benchmark instruments
08coupled bottlenecks
10open research programs
Working definition

A task is long-horizon when success depends on a sequence of interdependent decisions whose relevant state, cost, risk, or duration exceeds what the agent can safely solve as one local inference.

Dependency later steps rely on earlier state Persistence work crosses turns, windows, or sessions Exposure errors and side effects accumulate Uncertainty the environment reveals facts over time
How to read this track

Four layers, kept deliberately separate

The site distinguishes mechanisms, reported evidence, evaluation instruments, and research bets.

Knowledge map

Eight coupled bottlenecks

Click a mechanism to filter the source library. The boundaries are analytic, not architectural: a real system crosses all eight.

Unifying lens

The durable agent loop

Long-horizon competence emerges from the loop and its state architecture, not from any one prompt.

Attribution boundary

Do not confuse the model with the agent system

The same model can behave very differently under another interface, context policy, verifier, action budget, or sandbox.

observed capability = f(model, harness, tools, environment, budget, evaluator)
MODEL / INTERNAL

What travels with the weights

  • reasoning and world knowledge
  • tool-use policy and instruction following
  • long-context utilization
  • learned planning, recovery, and stopping behavior
  • trajectory policy improved by SFT / RL

Clean test: freeze the harness, tools, budget, and environment; swap only the model.

HARNESS / EXTERNAL

What wraps and routes the model

  • tool schemas, interfaces, permissions, and sandboxes
  • context selection, compaction, and persistent artifacts
  • decomposition, retries, checkpoints, and rollbacks
  • critics, tests, judges, and escalation policy
  • parallelism, resource ceilings, and stopping rules

Clean test: freeze the model and task set; ablate one harness component at a time.

CO-OPTIMIZATION

The boundary is porous

Agentic RL increasingly trains a model inside its deploy-time harness. A harness supplies the state and actions the model can learn; the trained model changes which scaffold components remain useful. Report both main effects and interactions with a crossed model × harness evaluation.

Evidence-backed synthesis

What the literature supports

Open each card for the source-level evidence, then inspect the synthesis, unresolved gap, and falsifiable hypotheses.

DIRECT reported in a paper OFFICIAL first-party engineering report SYNTHESIS inference across sources HYPOTHESIS testable, not established
Read the markdown claim ledger →
Measurement atlas

Benchmarks observe different horizons

A benchmark is an instrument. Its task unit, verifier, horizon, and hidden confounders determine what its score can mean.

Loading…
Source explorer

Search papers and official projects

Loading…

Evidence-derived research programs

Where the map is still weak

These are framed as measurable programs, not generic calls for “better planning” or “more memory.”

G1 · SCALING LAW

Reliability hazard by dependency depth

Estimate survival curves over consequential, dependent decisions; separate recoverable, irreversible, and silent failures.

Minimal test Hold model, harness, and domain fixed; procedurally vary dependency depth and checkpoint density.
G2 · CONTROL

Causal triggers for replanning

Represent assumptions and learn when new evidence invalidates only a subplan versus the global strategy.

Minimal test Inject observable, silent, and misleading blockers; measure repair cost and unnecessary plan churn.
G3 · STATE

Memory that knows when it is stale

Attach source, time, confidence, dependency, and expiry to memories; test update, contradiction, and selective forgetting.

Minimal test Use the same retrieved facts but vary whether the environment changed after storage.
G4 · OVERSIGHT

Verifier independence and placement

Measure correlated blind spots and allocate checks where expected risk reduction is highest.

Minimal test Cross generator, verifier, evidence access, and checkpoint placement under a fixed verification budget.
G5 · CONTEXT

Decision-aware compaction

Optimize summaries for future choices, not reconstruction or semantic similarity to the discarded transcript.

Minimal test Branch future tasks after compaction so relevance cannot be inferred from one known continuation.
G6 · INTERFACE

Action granularity and recoverability

Study when high-level tools reduce horizon and when they hide state, weaken diagnosis, or enlarge blast radius.

Minimal test Expose equivalent tasks through atomic, compositional, and macro tools with matched information.
G7 · LEARNING

Causal credit across the harness

Assign value to model calls, tool choices, memory writes, recoveries, and orchestration branches without shortcut rewards.

Minimal test Replay from stored checkpoints with counterfactual actions and compare credit estimators.
G8 · ATTRIBUTION

Crossed model × harness science

Measure main effects, interactions, transfer, complexity, and how quickly a scaffold's assumptions go stale.

Minimal test Evaluate several models under several frozen harnesses with identical tools, resources, and task seeds.
G9 · DEPLOYMENT

Pause, resume, handoff, and audit

Test agents that must stop safely, resume after external change, or transfer work to a fresh agent or human.

Minimal test Interrupt tasks at adversarial points; score resumption cost, state fidelity, and unsafe duplicate actions.
Open the full research program notes →
Reading paths

Enter from the system question

01

Why do agents fail with length?

tau-bench → METR Time Horizons → HORIZON → How Fast Do Agents Rot? → OSWorld 2.0

Lens: Does risk track duration, steps, dependencies, or irreversible state changes?
02

How should state survive?

Generative Agents → MemGPT → LongMemEval → MemoryAgentBench → Effective context engineering

Lens: What is authoritative state, what is experience, and what can be forgotten?
03

Where did the gain come from?

SWE-agent → Agentless → AI Agents That Matter → Agent Lightning v1.0 → Infrastructure noise

Lens: Model, interface, orchestration, resources, evaluator—or their interaction?
Research protocol

How this track handles evidence

1

Primary-first

Paper, project site, or official engineering report; secondary summaries do not anchor claims.

2

Scope-aware

Reported results stay attached to their model, harness, task, budget, and evaluator.

3

Claim-separated

Direct findings, official observations, synthesis, and hypotheses carry different labels.

4

Boundary-aware

Agent scores are treated as system measurements unless an experiment isolates the model.

Cutoff note. The library includes first-party sources available by September 24, 2026. Very recent preprints are marked adjacent or caveated when independent replication is limited.