Evidence-derived research programs
Where the map is still weak
These are framed as measurable programs, not generic calls for “better planning” or “more memory.”
G1 · SCALING LAW
Reliability hazard by dependency depth
Estimate survival curves over consequential, dependent decisions; separate recoverable, irreversible, and silent failures.
Minimal test
Hold model, harness, and domain fixed; procedurally vary dependency depth and checkpoint density.
G2 · CONTROL
Causal triggers for replanning
Represent assumptions and learn when new evidence invalidates only a subplan versus the global strategy.
Minimal test
Inject observable, silent, and misleading blockers; measure repair cost and unnecessary plan churn.
G3 · STATE
Memory that knows when it is stale
Attach source, time, confidence, dependency, and expiry to memories; test update, contradiction, and selective forgetting.
Minimal test
Use the same retrieved facts but vary whether the environment changed after storage.
G4 · OVERSIGHT
Verifier independence and placement
Measure correlated blind spots and allocate checks where expected risk reduction is highest.
Minimal test
Cross generator, verifier, evidence access, and checkpoint placement under a fixed verification budget.
G5 · CONTEXT
Decision-aware compaction
Optimize summaries for future choices, not reconstruction or semantic similarity to the discarded transcript.
Minimal test
Branch future tasks after compaction so relevance cannot be inferred from one known continuation.
G6 · INTERFACE
Action granularity and recoverability
Study when high-level tools reduce horizon and when they hide state, weaken diagnosis, or enlarge blast radius.
Minimal test
Expose equivalent tasks through atomic, compositional, and macro tools with matched information.
G7 · LEARNING
Causal credit across the harness
Assign value to model calls, tool choices, memory writes, recoveries, and orchestration branches without shortcut rewards.
Minimal test
Replay from stored checkpoints with counterfactual actions and compare credit estimators.
G8 · ATTRIBUTION
Crossed model × harness science
Measure main effects, interactions, transfer, complexity, and how quickly a scaffold's assumptions go stale.
Minimal test
Evaluate several models under several frozen harnesses with identical tools, resources, and task seeds.
G9 · DEPLOYMENT
Pause, resume, handoff, and audit
Test agents that must stop safely, resume after external change, or transfer work to a fresh agent or human.
Minimal test
Interrupt tasks at adversarial points; score resumption cost, state fidelity, and unsafe duplicate actions.