Attribution
Representation or score?
The headline comparison changes two things at once: the representation and the confidence
score. A 4×2 grid separates them, evaluating two scores over four representations — a frozen
multilingual encoder, a conventionally fine-tuned one, SetFit, and ours. The answer is not the one we expected.
The representation carries essentially the entire effect and the score carries almost none. Holding the
score fixed at MSP, moving from the CE baseline to the prototypical encoder takes RulingBR from 0.450 to 0.072;
holding the representation fixed, switching from MSP to our cosine score moves it from 0.072 to 0.073. Across
twelve corpora the mean absolute difference between the two scores over the prototypical representation is
0.0036 (median 0.0023), with no consistent winner. The accuracy-invariant panel says the same: over the
prototypical representation the two scores differ by at most 0.019.
Where the two representation-effect numbers on this page come from. Holding the score fixed at MSP and
changing only the geometry — from the CE encoder to the prototypical one, over the same backbone
— the median across the twelve corpora is 0.056 in E-AURC and 0.091 in correct-vs-incorrect
AUROC. Against 0.0023 and 0.014 on the score side, those give the ratios of 24 and 6.5 that bound the range. The
frozen-encoder contrast does not serve here, because it also swaps the base architecture. One identity is worth
stating, because we rely on it: since cos and MSP share the argmax, the error rate is the same under both
read-outs and the term E-AURC subtracts cancels in the difference, so Δ E-AURC = Δ AURC
exactly for the score contrast — which is why 0.0036 and 0.0023, measured in AURC, also hold in
E-AURC. The identity does not hold for the representation contrast, where accuracy changes, so the E-AURC column
on that side has to be computed.
It is worth bounding what that cell establishes. The two scores contrasted there are functions of the same
similarity vector over the same prototypes, so a small gap is partly expected by construction. The score-variant
appendix widens the set to five estimators, including X-Mahalanobis, which reads the same geometry through
cross-layer covariance and is the only one to stand apart — and even it sits 0.0096 from cosine, an order
of magnitude below the representation effect. All seven scores of the full grid have now
been evaluated over the prototypical representation, and the result separates three claims.
Energy and ReAct converge; the four structurally distinct ones do not. The Energy estimator demonstrates a median difference of +0.0003 relative to
cosine across all twelve corpora, while ReAct exhibits a +0.0001 margin, resulting in a spread of 0.0033 among the four estimators that leverage prototypical geometry
— comparable in magnitude to the 0.0023 margin of the cosine/MSP pairing. Conversely, the four structurally distinct scores remain behind:
kNN at +0.0439 from cosine at the median, ConjNorm at +0.0101, Mahalanobis at +0.0097 and GradNorm at +0.0069, each worse than cosine on eleven or
twelve of the twelve corpora. Given that all preserve the prototypical argmax output, these discrepancies constitute E-AURC differences, directly comparable to the 0.056
representation effect. When considering the spread across all eight estimators, the impact ratio between representation and score is reduced from 24 to approximately
1.2.
Adapting the representation halves the price of choosing badly. The same seven post-hoc scores show a median spread of 0.1060 AURC over the
cross-entropy encoder and 0.0460 over the prototypical representation — a 2.3-fold reduction. This is the attribution-ladder claim verified outside the family of
scores that motivated it, rather than only among readings of the same similarity vector.
Two observations bound this counterexample without invalidating it. The deficit is neither uniform nor a property of the latent representation: Mahalanobis and ConjNorm
actually beat cosine on RulingBR (0.071 against 0.073) and fall well behind on Buscape (0.136 and 0.121 against 0.057). What the four share is dependence on an input the
regime rations — a dense bank for kNN, an estimable covariance for Mahalanobis, a sampled partition constant for ConjNorm — and all must estimate it
from CK support examples. Consequently, the empirical evidence supports the narrow claim: over a task-adapted geometry, estimators provided with sufficient
operational conditions converge with one another and the cost of a bad choice halves — without the choice becoming indifferent.
From no adaptation to episodic prototypical training the gap falls from 0.107 to 0.014, a factor of about
eight, and the English replication gives 0.110 to 0.014 over its eight corpora. Which scalar is read off the
space matters most exactly where the space was never shaped for the task, and becomes almost irrelevant once it
was.
Two qualifications we state rather than round away. The step from a frozen encoder to a conventionally
fine-tuned one is not resolvable at this sample size (0.107 against 0.099); only the two steps involving
representation learning explicitly structured for the task — contrastive and episodic — survive, and
those are separated by a factor of four. And the ladder is monotone by median over the twelve Portuguese corpora
and by mean over the eight English ones, not corpus by corpus.
Decomposing the gap by score says why, and the answer is one-directional. This decomposition changes
scope, and we say so: it is the mean over the four deep-dive corpora of the 4×2 grid, not over the
twelve behind the medians above. Across the four rungs MSP rises 0.754 → 0.730 → 0.754 → 0.873 while the cosine
score rises 0.650 → 0.671 → 0.737 → 0.872. MSP is at
least as good as the cosine score on every representation in both languages; the gap closes because adaptation
rescues the geometric read-out until it merely ties. The honest statement is not that our confidence function is
better — it never is — but that on an adapted representation the choice of read-out stops costing
anything.
REFINEMENT 01
A frozen encoder: split verdict
On B2WReviews it attains the lowest AURC of the four representations (0.014) and certifies 0.98
coverage without a single gradient step. But on RulingBR it collapses to 0.498 — worse than a
conventionally fine-tuned BERTimbau and 6.8× behind ours — certifying no coverage at all. Pretrained
multilingual geometry suffices for generic high-resource tasks and is no substitute on specialist ones.
REFINEMENT 02
SetFit: membership in the family is not sufficient
Contrastive fine-tuning recovers part of the effect where the baseline is weakest (RulingBR
0.450 → 0.347) but not elsewhere, and even oracle-tuned stays 7× behind us on
B2WReviews and 4.5× on TuPy, certifying nothing on TuPy where we certify 0.23. How the
representation is restructured matters as much as whether it is.
REFINEMENT 03
At one shot, SetFit collapses entirely
At K = 1 it is worse than the CE baseline on three of four corpora and certifies
zero coverage everywhere, because its positive pairs degenerate into self-pairs when a class has a single
example — where an episodic prototype is still well defined, if noisy. We state the mechanism as a
hypothesis: we swept hyperparameters only at 5-shot.
This does not collapse into “fine-tuning yields a better classifier”
Correct-vs-incorrect AUROC is computed over ID inputs only and is invariant to how many the model gets right.
Against the Table 1 baseline it rises from 0.704 to 0.929 on RulingBR, 0.600 to 0.779 on TuPy, 0.811 to 0.917 on
B2WReviews and 0.806 to 0.863 on IntentPT. Matching accuracy by subsampling — discarding correct
predictions from the more accurate model until the two match, against the same swept baseline as the main tables
— gives AURC 0.029 vs. 0.048 on B2WReviews, 0.252 vs. 0.283 on IntentPT, 0.308 vs. 0.450 on RulingBR and
0.102 vs. 0.167 on TuPy — better on all four. Restricted to ID inputs: 0.029 vs. 0.048, 0.157 vs. 0.182,
0.206 vs. 0.340 and 0.102 vs. 0.167. The random discard is repeated 50 times per fold. The ranking advantage survives when the accuracy advantage
is removed entirely, on either population.
But the improvement is a property of the learned geometry, and attributing it to our choice of confidence
function — as an earlier framing of this work did — is not supported by the data.