STEERABILITY
RESEARCH SYSTEM / CONTROLLABLE BEHAVIOR

Can we move a model where we intend—and nowhere else?

Steerability is the ability to move a model's behavior toward a specified target with the right magnitude, low off-target change, and reliable transfer across inputs and environments. This map connects prompting, activation interventions, preference learning, model editing, evaluation, and scalable oversight.

Living research map Reviewed 32 representative papers · claims separated from hypotheses
Core distinction
capability: ∃ an output the model can produce ≠ steerability: a user can reliably reach it
Operational definition

Steering is a control problem, not a vibe.

A useful definition has to specify the target, the intervention channel, the behavioral measurement, the budget, and the distribution over contexts.

target displacementΔz*
→
interventionu
→
model + contextF(x, u)
→
observed displacementΔẑ
01

Reachability

Can a target be elicited through an available interface, not merely sampled somewhere in the model's support?

target ∈ reachable set?
02

Calibration

Does changing intervention strength produce the requested amount of behavioral movement without under- or overshooting?

gain = ‖Δẑ∥ / ‖Δz*‖
03

Selectivity

How much movement leaks into dimensions the user did not ask to change?

leakage = ‖Δẑ⊥‖
04

Robustness

Does the same controller work across paraphrases, tasks, users, model versions, adversaries, and distribution shift?

worst-case, not only mean

Scope. Instruction following is one observable slice of steerability. Alignment is broader: it asks whether the chosen target is the right one and whether the system remains aligned when direct steering is weak, ambiguous, or adversarial.

Operational taxonomy

Nine lenses, one closed loop

The categories separate what is controlled, where the intervention enters, and how success and failure are measured. Click a lens to filter the library.

Interactive conceptual model

A controller can hit the axis and still miss the goal.

Adjust strength and entanglement. Strength changes movement along the requested axis; entanglement creates an unintended side effect. This toy diagram is explanatory—not a reported model result.

2D GOAL-SPACE conceptual
requested attribute off-target attribute baseline target response
target gain0.80×
side effect0.24
total error0.31
Evidence-backed synthesis

What survives across method families

Each claim distinguishes reported evidence from synthesis. The details and paper-level limitations remain visible in the library.

01
STRONG SYNTHESIS

Capability is not reachability.

A model can contain or occasionally sample a behavior that users cannot reliably elicit. Steerability therefore depends on the reachable set induced by the interface, budget, and user—not only the model's support.

EvidenceVafa et al. ↗Chang et al. ↗Human reproduction tasks and uniformly sampled goal-space probes both expose gaps hidden by capability metrics.
02
STRONG SYNTHESIS

Scalar success hides vector error.

Following the requested dimension can coexist with drift in style, refusal, truthfulness, helpfulness, or other latent goals. A controller should be scored by its full response vector and cost.

EvidenceCourse Correction ↗ITI ↗SAE refusal ↗Side effects persist under prompting and RL; truthfulness and safety interventions expose explicit trade-offs.
03
MODERATE SYNTHESIS

Linear handles are real—and dangerously easy to overread.

Low-dimensional interventions can causally alter behavior, but linear efficacy does not prove that the concept itself is one-dimensional, mechanistically localized, or stable across contexts.

EvidenceCAA ↗Refusal direction ↗AxBench ↗Activation directions work, yet simple prompts can outperform tested representation methods and low-rank safety handles can be removed.
04
STRONG SYNTHESIS

Training improves the default; conditioning preserves plurality.

RLHF and DPO move a model toward an aggregate preference distribution. Runtime control requires the objective itself to remain conditional on user, constitution, authority, or explicit attributes.

EvidenceInstructGPT ↗SteerLM ↗Prompt steerability ↗Post-training strengthens instruction following, while explicit attribute conditioning and persona curves reveal what a single global policy loses.
05
STRONG SYNTHESIS

Robust control needs authority and provenance, not better wording alone.

When trusted instructions and untrusted data share one token stream, the model must infer which source has authority. Training or architecture must represent that boundary explicitly.

EvidenceInstruction Hierarchy ↗StruQ ↗Sleeper Agents ↗Hierarchy and channel separation help; conditional policies can nevertheless survive ordinary safety training.
06
RESEARCH HYPOTHESIS

Oversight quality should be treated as closed-loop steerability.

A weak supervisor does not merely label a strong model; it perturbs a learning system and observes a partial response. The key object is the loop gain from weak evidence to strong behavior under misspecification.

Evidence baseWeak-to-strong ↗Weak judges ↗Reward tampering ↗Strong students exceed weak teachers but do not recover full capability; oversight protocols are task-dependent and proxies can induce harmful generalization.
Adjacent research programs

Where steerability becomes alignment research

These links are not synonyms. They specify which variable each neighboring field contributes to the control loop.

ProgramContribution to steerabilityFailure exposedJoint research question
Weak-to-strong

Weak supervision supplies a noisy, low-bandwidth steering signal to a more capable learner.

Imitation of weak errors or overwriting latent strong knowledge.

When does directional supervision elicit a strong feature instead of collapsing to the weak policy?

Alignment

Chooses or learns the target; steerability measures whether an intervention reaches it.

A system can be precisely steerable toward the wrong objective—or aligned on average but rigid for individual users.

How should target legitimacy, plurality, and controllability be evaluated jointly?

Control theory

Contributes reachability, gain, coupling, stability, feedback, disturbance rejection, and robust control.

Open-loop success at one prompt says little about closed-loop behavior under drift or adversaries.

What are the controllability Gramian, condition number, and stability margins of a generative model in behavior space?

Model organisms

Create controlled failure modes whose causes are known, enabling causal tests of detectors and counter-steering.

Mitigations may solve the organism while missing natural failures; realism and controllability trade off.

Which organisms predict mitigation performance on naturally emerging misalignment?

Scalable oversight

Builds feedback channels that let weak judges steer stronger systems on tasks they cannot solve unaided.

Persuasion, shared blind spots, evaluator gaming, and information asymmetry.

Can debate, critique, or decomposition improve both judge accuracy and the resulting policy's robust steerability?

Kaijing research agenda

Six projects that can become experiments

Each program starts from an empirical contradiction or limitation, proposes a minimal study, and has a result that could falsify the motivating hypothesis.

01 · FOUNDATIONS

Cross-method steering response

Question. How comparable are the behavioral effects produced by prompt, activation, and preference-based steering?

Minimal test. Use one contrast dataset to derive all three interventions; fit a local Jacobian from intervention coefficients to a shared multi-dimensional goal-space.

Falsifier. Cross-method directions remain unrelated after controlling target, model, layer, scale, and evaluation.

02 · SUPERALIGNMENT

Weak-to-strong steerability transfer

Question. Can a weak supervisor specify a direction of improvement without specifying the full strong policy?

Minimal test. Compare weak labels, weak pairwise differences, weak critiques, and weak activation contrasts at equal annotation budgets.

Falsifier. Directional supervision never recovers more strong capability than ordinary weak-label fine-tuning.

03 · EVALUATION

Controllability under shift

Question. Which interventions preserve calibrated gain and low coupling across paraphrases, domains, languages, and model updates?

Minimal test. Estimate a response matrix in-distribution, then predict the full output displacement OOD without retuning.

Falsifier. No method beats per-domain prompt tuning once total adaptation cost is counted.

04 · REPRESENTATIONS

Adaptive low-rank control

Question. Can a controller choose layer, token, direction, and strength per input while obeying a side-effect budget?

Minimal test. Train a sparse router against target gain, orthogonal leakage, KL change, and latency; compare fixed CAA, SAE, prompt, and LoRA controls.

Falsifier. Adaptive routing improves the target only by moving hidden side effects outside the measured axes.

05 · MODEL ORGANISMS

Counter-steering benchmarks

Question. Does a mitigation remove a misaligned policy or merely suppress its current readout?

Minimal test. Apply prompt, activation, preference, and edit-based counter-steering to sleeper, reward-gaming, and emergent-misalignment organisms; then change triggers and oversight.

Falsifier. Mitigation success on controlled organisms has no predictive value for new triggers or newly trained organisms.

06 · OVERSIGHT

Judge-aware closed-loop control

Question. Can oversight adaptively choose critique, debate, decomposition, or abstention based on its estimated control authority?

Minimal test. Give a weak judge a fixed budget and actions with known costs; optimize downstream strong-model improvement rather than judge accuracy alone.

Falsifier. Better judge decisions fail to produce better or more robustly steerable learned policies.

Paper knowledge base

Search claims, experiments, and limitations

Loading…

Core steerability evidenceMethodological rootAlignment bridgeExpand any card for experiments, limitations, and future work.
Evidence standard

A map should expose what it does not know.

DIRECT Results and limitations reported by the cited paper.

SYNTHESIS Cross-paper inference with the supporting chain visible.

HYPOTHESIS A falsifiable proposal, including a result that would count against it.