LOW-PRECISION RL GitHub ↗
Living research map

When the policy is
numerically approximate,
what actually changes?

Low precision enters a feedback loop. It can accelerate rollout, reshape exploration, move the behavior policy, distort importance ratios, erase small updates—and decide which gradients survive.

8core dossiers
12indexed claims
5evidence labels
8research programs

Precision is structural,
not uniform.

The emerging evidence does not say “FP8 works” or “quantization is unstable.” It says ordinary dense computation, persistent state, routing decisions, probability ratios, clipping boundaries, and gradients have different numerical needs.

Research question A

When does numerical error become useful stochasticity versus destructive policy mismatch?

Research question B

Which parts of policy optimization actually require high numerical precision?

A small error gets several chances to become a policy change.

↺ The learner changes the next behavior policy. Numerical error is therefore part of the data-generating process, not a one-time inference perturbation.

Do not average away the disagreement.

The conflict often reveals the missing variable.

QeRLvsQaRL

Entropy: exploration or failure?

QeRL reports useful entropy from NVFP4/AQN in LoRA RL. QaRL sees no general elevation and links late entropy growth to repetitive error tokens.

Boundary candidates · update regime · noise schedule · objective · horizon
Jet-RL→Full FP8

Alignment is necessary—not sufficient.

A unified FP8 path reduces mixed-precision drift. Later work shows the ratio of two FP8 probabilities can still corrupt the trust region.

Missing variable · objective semantics under noisy ratios
FP8-RLvsQaRL

Does FP8 KV cache pay off?

One 20K, preemption-heavy setup reports 38% from KV alone. Another vLLM setup reports no meaningful throughput gain.

Boundary candidates · backend · concurrency · preemption · length

Locate a method by what it perturbs.

Select a program to filter the paper map.

Papers as evidence,
not trophies.

Every entry records regime, result, limitation, and numerical path.

Showing all research programs

What is known—and how?

Filter by epistemic status. Conflicts remain visible.

Read the full source-level ledger →

The gradient that vanished too early.

One failure mechanism in full-pipeline FP8 connects tiny probability errors to macroscopic entropy collapse.

See all eight failure modes →
  1. Small FP8 probability errorscurrent and old policy forwards
  2. Ratio error compoundsdivision amplifies both errors
  3. Negative tokens over-clipthe lower trust boundary moves
  4. Corrective gradients vanishgarbled trajectories are not suppressed
  5. Pathology proliferatesentropy surge · reward loss · collapse

Questions that can become experiments.

Each card contains a discriminating test and a falsifiable prediction.

Enter by the problem you want to solve.