Reachability
Can a target be elicited through an available interface, not merely sampled somewhere in the model's support?
target ∈ reachable set?
Steerability is the ability to move a model's behavior toward a specified target with the right magnitude, low off-target change, and reliable transfer across inputs and environments. This map connects prompting, activation interventions, preference learning, model editing, evaluation, and scalable oversight.
capability: ∃ an output the model can produce
≠
steerability: a user can reliably reach it
A useful definition has to specify the target, the intervention channel, the behavioral measurement, the budget, and the distribution over contexts.
Can a target be elicited through an available interface, not merely sampled somewhere in the model's support?
target ∈ reachable set?
Does changing intervention strength produce the requested amount of behavioral movement without under- or overshooting?
gain = ‖Δẑ∥ / ‖Δz*‖
How much movement leaks into dimensions the user did not ask to change?
leakage = ‖Δẑ⊥‖
Does the same controller work across paraphrases, tasks, users, model versions, adversaries, and distribution shift?
worst-case, not only mean
Scope. Instruction following is one observable slice of steerability. Alignment is broader: it asks whether the chosen target is the right one and whether the system remains aligned when direct steering is weak, ambiguous, or adversarial.
The categories separate what is controlled, where the intervention enters, and how success and failure are measured. Click a lens to filter the library.
Adjust strength and entanglement. Strength changes movement along the requested axis; entanglement creates an unintended side effect. This toy diagram is explanatory—not a reported model result.
Each claim distinguishes reported evidence from synthesis. The details and paper-level limitations remain visible in the library.
A model can contain or occasionally sample a behavior that users cannot reliably elicit. Steerability therefore depends on the reachable set induced by the interface, budget, and user—not only the model's support.
Following the requested dimension can coexist with drift in style, refusal, truthfulness, helpfulness, or other latent goals. A controller should be scored by its full response vector and cost.
Low-dimensional interventions can causally alter behavior, but linear efficacy does not prove that the concept itself is one-dimensional, mechanistically localized, or stable across contexts.
RLHF and DPO move a model toward an aggregate preference distribution. Runtime control requires the objective itself to remain conditional on user, constitution, authority, or explicit attributes.
When trusted instructions and untrusted data share one token stream, the model must infer which source has authority. Training or architecture must represent that boundary explicitly.
A weak supervisor does not merely label a strong model; it perturbs a learning system and observes a partial response. The key object is the loop gain from weak evidence to strong behavior under misspecification.
These links are not synonyms. They specify which variable each neighboring field contributes to the control loop.
Weak supervision supplies a noisy, low-bandwidth steering signal to a more capable learner.
Imitation of weak errors or overwriting latent strong knowledge.
When does directional supervision elicit a strong feature instead of collapsing to the weak policy?
Chooses or learns the target; steerability measures whether an intervention reaches it.
A system can be precisely steerable toward the wrong objective—or aligned on average but rigid for individual users.
How should target legitimacy, plurality, and controllability be evaluated jointly?
Contributes reachability, gain, coupling, stability, feedback, disturbance rejection, and robust control.
Open-loop success at one prompt says little about closed-loop behavior under drift or adversaries.
What are the controllability Gramian, condition number, and stability margins of a generative model in behavior space?
Create controlled failure modes whose causes are known, enabling causal tests of detectors and counter-steering.
Mitigations may solve the organism while missing natural failures; realism and controllability trade off.
Which organisms predict mitigation performance on naturally emerging misalignment?
Builds feedback channels that let weak judges steer stronger systems on tasks they cannot solve unaided.
Persuasion, shared blind spots, evaluator gaming, and information asymmetry.
Can debate, critique, or decomposition improve both judge accuracy and the resulting policy's robust steerability?
Each program starts from an empirical contradiction or limitation, proposes a minimal study, and has a result that could falsify the motivating hypothesis.
Question. How comparable are the behavioral effects produced by prompt, activation, and preference-based steering?
Minimal test. Use one contrast dataset to derive all three interventions; fit a local Jacobian from intervention coefficients to a shared multi-dimensional goal-space.
Falsifier. Cross-method directions remain unrelated after controlling target, model, layer, scale, and evaluation.
Question. Can a weak supervisor specify a direction of improvement without specifying the full strong policy?
Minimal test. Compare weak labels, weak pairwise differences, weak critiques, and weak activation contrasts at equal annotation budgets.
Falsifier. Directional supervision never recovers more strong capability than ordinary weak-label fine-tuning.
Question. Which interventions preserve calibrated gain and low coupling across paraphrases, domains, languages, and model updates?
Minimal test. Estimate a response matrix in-distribution, then predict the full output displacement OOD without retuning.
Falsifier. No method beats per-domain prompt tuning once total adaptation cost is counted.
Question. Can a controller choose layer, token, direction, and strength per input while obeying a side-effect budget?
Minimal test. Train a sparse router against target gain, orthogonal leakage, KL change, and latency; compare fixed CAA, SAE, prompt, and LoRA controls.
Falsifier. Adaptive routing improves the target only by moving hidden side effects outside the measured axes.
Question. Does a mitigation remove a misaligned policy or merely suppress its current readout?
Minimal test. Apply prompt, activation, preference, and edit-based counter-steering to sleeper, reward-gaming, and emergent-misalignment organisms; then change triggers and oversight.
Falsifier. Mitigation success on controlled organisms has no predictive value for new triggers or newly trained organisms.
Question. Can oversight adaptively choose critique, debate, decomposition, or abstention based on its estimated control authority?
Minimal test. Give a weak judge a fixed budget and actions with known costs; optimize downstream strong-model improvement rather than judge accuracy alone.
Falsifier. Better judge decisions fail to produce better or more robustly steerable learned policies.
Loading…
DIRECT Results and limitations reported by the cited paper.
SYNTHESIS Cross-paper inference with the supporting chain visible.
HYPOTHESIS A falsifiable proposal, including a result that would count against it.