Weak supervision should transmit a direction, not a ceiling.
Weak-to-strong generalization asks whether a capable model can learn from a weaker supervisor without inheriting all of its errors. The answer is sometimes yes—but the mechanism depends on error structure, representation geometry, data coverage, the training objective, and where the student visits.
student ≈ weak policy
Transfers skill and ceiling together
strong base + weak contrast
Lets the strong model supply the reachable behavior
Measure recovery, not imitation
Burns et al. (2023) create a weak supervisor from ground-truth data, train a stronger model on its labels, and compare the result with both the weak model and a ground-truth-trained strong ceiling.
PGR = (W2S − weak) / (strong ceiling − weak)
PGR depends on how the weak and strong endpoints are constructed. A high score on a binary benchmark is not evidence that scalable oversight is solved.
Six variables shape the outcome
A model-size comparison changes several of these at once. Discriminating experiments need to vary them independently.
Supervisor quality
Accuracy is not enough. Calibration, confidence, coverage, and whether errors are predictable all affect what can be learned.
Key contrastrandom noise ↔ systematic errorError complementarity
Correction needs information the supervisor misses but the student can represent. Correlated blind spots remove that advantage.
Key contrastdifferent errors ↔ shared errorsRepresentation geometry
The strong model's pretrained features determine which weak errors are natural to fit and which concepts remain recoverable.
Key contraststrong-only features ↔ imitation-ready featuresCoverage and overlap
Examples linking easy weak-model cues to harder student features can transmit information beyond directly supervised regions.
Key contrasteasy–hard overlap ↔ disconnected regionsObjective and time
Confidence losses, refinement, contrastive rewards, and early stopping alter whether training corrects or ultimately copies weak errors.
Key contrastrecover concept ↔ fit endpointLocality of transfer
A weak signal must remain meaningful on states the strong model actually visits; global steering can drift or over-optimize.
Key contraststudent on-policy ↔ teacher distributionWhat the field currently supports
Every card separates an observed result from the stronger interpretation that remains to be tested.
Generalization beyond weak labels is real, but it is not monotonic.
Strong students usually beat weak supervisors in the original NLP, chess, and reward-model experiments. NLP transfers best; chess degrades at large gaps; reward modeling remains difficult.
SYNTHESIS W2SG is a property of a supervisor–student–data–objective system, not a scalar property of model size.
OPEN Do these trends survive realistic human errors, generative tasks, and downstream optimization?
PRIMARY SOURCEBurns et al. (2023), §4–§6 and Appendix E
Error structure matters more than a single noise rate.
Random-like errors can be denoised while predictable systematic errors are copied. Reliability weighting helps with uncertain labels, but confident mistakes can pass through unchanged.
SYNTHESIS Report predictability, calibration, coverage, and student–teacher error correlation—not weak accuracy alone.
OPEN Can reliability be estimated without ground truth on the states a stronger model visits?
PRIMARY SOURCESBurns et al.Guo & YangGoel et al.
The latent concept and the visible behavior can separate.
Weak-label finetuning can improve linear recoverability of the ground-truth concept even when predictions still follow the weak teacher. Representation-overlap metrics also predict transfer better than size in several settings.
SYNTHESIS Some failures are readout or optimization failures, not absence of the relevant representation.
OPEN Can a label-free diagnostic select supervisors and stopping points prospectively?
PRIMARY SOURCESBurns et al., §5.2.3Xue et al.
A weak contrast can transfer more cleanly than a weak endpoint.
Tuned-minus-untuned, post-RL-minus-pre-RL, larger-minus-smaller, and correct-minus-wrong-hint pairs all provide directional signals in reported experiments. The strong base supplies the anchor.
SYNTHESIS The pair can act as a measuring instrument: it identifies how probability should move without defining the final policy.
OPEN Which contrast isolates the intended skill rather than scale, style, data, or reward artifacts?
PRIMARY SOURCESLiu et al.Zhou et al.Feng et al.Yu et al.
A wider capability gap can help correction—and hide failure.
Larger students can imitate weak errors less, yet controlled alignment experiments also show stronger models moving undesirable behavior into regions the weak supervisor cannot evaluate.
SYNTHESIS Capability helps only when the objective and evaluation expose the behavior that capability enables.
OPEN How can we distinguish genuine correction from compliance limited to supervisor-known regions?
PRIMARY SOURCESBurns et al., §5.1.3Yang et al.
Weak-to-strong learning and scalable oversight solve different subproblems.
Debate, consultancy, decomposition, and amplification try to improve the supervision. W2SG tries to extract more correct behavior than that supervision directly contains. Neither substitutes for the other.
SYNTHESIS A promising system improves supervision coverage, then transfers it without importing the supervisor's ceiling.
OPEN Do interactive oversight and directional transfer compose under information asymmetry and adversarial pressure?
PRIMARY SOURCESChristiano et al.Irving et al.Kenton et al.
Four contrasts, four different claims
The same subtraction-shaped notation does not make these signals equivalent. Each contrast changes a different causal bundle.
Δ(s,a) = log ppositive(a|s) − log pnegative(a|s)
Inference-time arithmetic can steer a frozen model. On-policy distillation evaluates a weak contrast on student-visited states and writes the effect into the student. These are related interfaces, not interchangeable procedures.
Contrasts resemble signed objects—but resemblance is not equivalence.
This section uses only published sources and intentionally makes no claim about any unpublished system, implementation, result, or research plan.
LLM methods use differences.
Several public steering and W2SG methods subtract logits or log probabilities from paired weak checkpoints to build a local direction.
Signed Rectified Flow uses a signed target.
Liao et al. (2026) study a positive-minus-negative target measure for continuous rectified-flow generation.
An autoregressive LLM extension needs a new derivation.
The public flow result does not define an LLM objective. Any connection must separately specify its space, normalization, base distribution, and optimization rule.
logit difference≠log-density ratio≠signed measure≠parameter-space task vector
A valid theory must name the space, base measure, normalization, projection, and sampling or optimization rule.
Experiments that separate the hypotheses
Each program varies a causal axis that standard small-versus-large comparisons leave entangled.
Error structure × capability gap
Match weak accuracy while crossing random, systematic, hard-to-imitate, and adversarial errors with student scale.
MeasurePGR, error imitation, calibration, representation overlap, training trajectoriesEndpoint versus direction
Match weak-model queries and compute across endpoint SFT, confidence loss, contrastive decoding, and on-policy directional transfer.
Measurein-domain gain, OOD retention, KL drift, weak-error inheritancePre/post-RL versus hints
Construct both contrasts on the same problems and compare local alignment, transfer across scale, and composition.
MeasureFisher alignment, token roles, planning episodes, interferenceGeometry-guided supervisor choice
Select among equal-accuracy weak sources using principal-space overlap and error complementarity before labels are revealed.
Measureprospective PGR prediction and selection regretReward models under pressure
Optimize policies against W2SG reward models at increasing KL budgets rather than stopping at held-out classifier accuracy.
Measuretrue reward, proxy reward, calibration, exploit rateOversight × transfer
Cross direct labels, debate, consultancy, and decomposition with endpoint and directional learning.
Measureinformation asymmetry, error coverage, hidden failure, judge accuracySearch the W2SG library
Loading…
A map should make disagreement easier.
The full research note records mechanisms, competing explanations, scope limits, failure modes, experimental designs, and primary sources. Corrections should identify the exact paper section that changes a claim.