Sensitivity or Criterion?Measuring and Moving When Language Models Act
Paper arXiv soon Code and data soon BibTeX
Two models can have the same accuracy and need opposite fixes.
-
One accuracy hides two failures. Before an agent acts it must decide whether to act at all. A model that fails to see when acting is warranted and one that sees it and holds back get the same score.
-
FlipAct measures them separately. A sensitivity and a criterion for every model, on an evidence scale computed by a rule, so the numbers are comparable across models.
-
The criterion can be moved on its own. Most open models hold back more than the evidence warrants. Steering and distillation shift that, with no significant change in sensitivity.
Why one accuracy is not enough
Three short steps from psychology: what an accuracy hides, how signal-detection theory splits it, and where the same split shows up in agents.
Two observers, one accuracy, opposite mistakes
Start with a case where the stakes are obvious. Two doctors each read the scans of 100 patients who have a tumour and 100 who are healthy. Their accuracy is 65% and 64%. By that number they are the same doctor.
Accuracy adds the two kinds of error together. What to do about them depends on which kind you have.
Sensitivity and criterion: try it on yourself
Signal-detection theory was worked out in the 1950s for radar operators and for listeners in hearing experiments. It says every yes-or-no judgement under uncertainty comes from two separate things.
Sensitivity d′
How well you can tell a signal from none. It is the distance between the two curves.
Criterion c
How sure you want to be before you say yes. It is where you put the line.
Make a miss expensive and people move the line left. Make a false alarm expensive and they move it right. What they can see has not changed. Try it: a patch of static will flash, and half the time it hides a faint blob.
- You answer yes or no, and see right away whether you were right.
- After 20 trials you get two numbers: how well you can tell, and how freely you say yes.
Now the full picture, with both knobs. It is the same for a listener, for the two doctors above, and for a language model deciding whether to call a tool.
Moving the line trades one error for the other. Only pulling the curves apart removes errors.
The same split, wherever an agent decides to act
The doctors' problem is every agent's problem: act, or hold back. Pick a decision and a kind of model, and see which error it makes and what that error costs.
FlipAct
What the paper builds and what it finds: an evidence scale computed by rule, a sensitivity and a criterion for 44 open models, and two ways to move the criterion alone.
Evidence computed by rule, not written by a model
To compare criteria across models, every question needs an evidence score on one scale with a fixed zero. FlipAct computes it from qualitative physics: directions and coarse step sizes, no physical units.
A state
Slightly, very or dangerously off, in one of two directions. Normal is zero.
A tool
Pushes the state one way, by one, two or three steps.
A score, ΔU
How much closer to normal the tool leaves the state. Yes exactly when ΔU > 0.
The design unit is a 2 × 2 quad: two tools, each shown in the state it fixes and in the state it worsens. Every question of a quad needs the same two facts, what the tool does and what the room needs. Only the state changes, and with it the answer, so a model that knows what a humidifier is for must still read the state.
Language models help only in the two indigo steps. The rule decides every state, answer and score.
Reading a model: a slope and a crossing point
Because every question carries a score, a model's Yes-rate can be fitted as a curve over that score. The slope a is its sensitivity, in the role of d′. The score τ where the curve crosses one half is its criterion. The rule's boundary sits at zero for every model, so a positive τ means the model says Yes less often than the rule warrants.
Sensitivity behaves like a capability
In the Qwen3 family, bigger models tell helpful from harmful actions better. The ranking agrees with sensitivity on other benchmarks.
The criterion is conservative by default
Most open models hold back more than the evidence warrants.
It is a per-model setting, not a size effect
Two releases of Mistral-Small-24B, the same size, land on opposite sides of zero.
Models read the sign of the evidence, not its size. Actions that leave the state no better are accepted about as often as helpful ones.
Moving the criterion, and only the criterion
If the criterion is a separate setting, it should move while sensitivity stays put. Three interventions, two of which aim at exactly that.
Steering
At inference. A direction read from the model's own activations is added during generation. Weights stay fixed.
moves the criterionCriterion distillation
Into the weights. The model is distilled from its own steered copy.
moves the criterionSensitivity learning
Training on the computed labels does the opposite job. The curve gets steeper.
raises sensitivityIt carries to other benchmarks
After distillation, the criterion moves toward acting on every scored model and benchmark, across six external benchmarks.
Sensitivity is kept
On the external benchmarks, how well the model tells the cases apart stays within this.
Accuracy follows the starting point
Models that start conservative gain accuracy. Models that start permissive do not.
Steering, with the weights fixed
On a tool-calling benchmark, steering moves all five models toward acting.
Show the numbers
| call rate | criterion c | ||||
|---|---|---|---|---|---|
| Model | unsteered | steered | unsteered | steered | Δ accuracy |
| Qwen3-8B | 0.56 | 0.60 | −0.18 | −0.32 | −0.022 |
| Qwen3.5-27B | 0.47 | 0.50 | +0.10 | +0.01 | +0.005 |
| gemma-2-9b | 0.24 | 0.29 | +0.78 | +0.59 | +0.002 |
| Mistral-24B | 0.37 | 0.55 | +0.39 | −0.13 | −0.032 |
| phi-4 | 0.57 | 0.59 | −0.30 | −0.33 | −0.038 |
A when-to-act benchmark says more when it reports d′ and c beside accuracy.
BibTeX
@article{dong2026flipact,
title = {Sensitivity or Criterion? Measuring and Moving When Language Models Act},
author = {Dong, Sixun and Hu, Yebowen and Yin, Ming and Chen, Yiran and Liu, Fei and Chen, Chen},
year = {2026}
}
Questions or collaboration: [email protected]