Long paper · under review

Selective Classification in Few-Shot Text: The Effect of Representation on Confidence Heuristics

12 / 12corpora where prototypical geometry beats every post-hoc baseline at 5-shot
−67%median AURC reduction against the strongest baseline under an oracle sweep
24×how much more the representation moves performance than a score reading the same geometry, in E-AURC (median 0.056 against 0.0023); across all eight scores evaluated, 1.2×
0.91certified coverage at 5% risk on court rulings, under calibration independent of model selection, where the tuned baseline is infeasible
15.2%risk actually realised by that 5%-certified threshold once unseen classes enter the stream
Overview

Where does the gain actually come from?

This is a measurement and attribution study, not a proposal of new model components. Few-shot text classifiers abstain poorly, and the field has responded by designing better confidence scores over a representation taken as given. We evaluate two scores over four representations — and then eight scores over the prototypical representation — and find the effect is carried almost entirely by the representation.

Abstract

Few-shot text classifiers rank their own errors badly: under severe data scarcity, confidence estimates do not separate correct from incorrect predictions well enough to sustain useful coverage under a risk constraint. We evaluate twelve Portuguese corpora at budgets of 1 to 10 examples per class, against seven post-hoc baselines — the strongest of them under an oracle hyperparameter sweep. Episodic prototypical fine-tuning shifts that frontier: AURC falls by a median of 67% at 5-shot. On a calibration set disjoint from model selection as well, the method certifies 0.91 coverage at a 5% target risk on court rulings, where the tuned baseline admits no feasible threshold at all.

A 4×2 factorial experiment separates representation from confidence estimator, and locates the effect in the representation. Measured on quantities that discount accuracy, changing the representation moves performance 6 to 24 times more than changing the score, depending on the quantity: a median of 0.056 against 0.0023 in E-AURC and of 0.091 against 0.014 in correct-vs-incorrect AUROC, across the twelve corpora. That range is measured against the estimators that read the prototypical geometry: evaluating all seven post-hoc scores of the grid there, four of them — denied by the regime the input they need — fall behind and narrow the ratio to about 1.2. Two caveats qualify the finding. Representation interventions are not interchangeable — SetFit, even at the configuration chosen by its own sweep on validation AURC, trails well behind and degrades severely at one example per class. And the certificate bounds risk only in-distribution: in a stream containing unseen classes, the same threshold reaches 15.2% empirical error against the 5% target, on the worst of the five corpora evaluated.

CONTRIBUTION 01

The broadest few-shot selective-classification evaluation in Portuguese

Twelve corpora, five task families, three budgets, five draws each. At 5-shot, prototypical geometry achieves lower AURC than every post-hoc baseline on every corpus, cutting it by a median of 67% against the strongest of them under an oracle hyperparameter search.

CONTRIBUTION 02

A quantitative attribution: representation vs. score

A 4×2 grid showing the score contributes less than 1% of the representation's effect — and that this dominance is a property of the few-shot regime, not of any component we introduce, since the ablation finds the label adapter and QDA sampler both dispensable.

CONTRIBUTION 03

Certified coverage on all twelve corpora

With the calibration size m reported per corpus, and the realised risk once unseen classes enter the stream — three times the target, and the number a deployment decision actually turns on.

CONTRIBUTION 04

Cross-lingual replication and two controls

The attribution replicates on eight English corpora, the ladder's extremes within 0.003. A second backbone (XLM-R) preserves the advantage on nine of twelve corpora but not its magnitude, and three risk-control rules are compared over our own representation.

SCOPE

What we adopt rather than propose

SGR is an existing post-hoc framework (Geifman & El-Yaniv, 2017) and LAQDA an existing backbone whose two mechanisms we find dispensable. The operative ingredient is a plain prototypical encoder, available to any practitioner already using prototypical networks. We call the system evaluated here ProtoSel — episodic prototypical fine-tuning, the geometric score and the SGR threshold together — to keep it distinct from LAQDA, which is only the inherited backbone.

PRIOR WORK

The ordering is known in vision; the magnitude is not

That representation quality dominates the post-hoc estimator is established for data-rich ImageNet classifiers. Our contribution is its magnitude in the few-shot regime. Measured on quantities that discount accuracy, it lands between 6 and 24 times, depending on the quantity: a median of 0.056 against 0.0023 in E-AURC and of 0.091 against 0.014 in correct-vs-incorrect AUROC. Sensitivity to the score falls roughly eightfold with adaptation.

Why it matters

A clarification paper, not another score

A paper that said “we invented a cosine score that improves AURC by 5%” would be one more entry in a crowded field. This one says: we tested seven modern scores and none of them is what matters; what solves selective risk under few-shot supervision is episodic prototypical geometry — and we isolated the variables to show it.

01

It closes a direction the field is still pouring effort into

Post-hoc OOD detection and selective classification have advanced largely by designing ever more elaborate confidence functions — Mahalanobis, kNN, GradNorm, ReAct, ConjNorm — over a representation taken as given. Our 4×2 attribution shows that in the few-shot regime this axis is close to exhausted. Measured on quantities that discount accuracy — the honest comparison, since swapping the score holds accuracy fixed by construction and swapping the representation does not — the representation moves performance 6 to 24 times more than the score, depending on the quantity: a median of 0.056 against 0.0023 in E-AURC and of 0.091 against 0.014 in correct-vs-incorrect AUROC, across twelve corpora. In raw AURC the ratio exceeds 100, but part of that is arithmetic. Sensitivity to the score is itself a function of adaptation, falling roughly eightfold from a frozen encoder to an episodically trained one. The finding cuts against our own earlier framing, which is what makes it worth reporting: the empirical effort should move from post-hoc scores to the structure of the latent space.

02

It decides whether a risk-controlled system is deployable at all

Nobody had shown whether a distribution-free rule such as SGR could lift coverage off zero when the representation itself is estimated from five examples per class. Under a disjoint calibration split the tuned baseline admits no feasible threshold at 5% risk on six of the twelve corpora; our selector certifies non-zero coverage on four of those six. On RulingBR the tightest bound MSP reaches by scanning every threshold is b* = 0.211 — and even with no correction at all for the multiplicity of the search, the floor stays at 0.149, well above the 5% target. Our selector certifies 0.94 coverage there (0.91 under calibration independent of model selection). That is the difference between a system that can be deployed under a risk constraint and one that cannot.

03

Scale is not a substitute for geometry

A prompted 7 B model is more accurate than our 110 M pipeline on reviews (0.973 vs. 0.971) and still ranks its own errors worse (AURC 0.011 vs. 0.004), at roughly 2,756 s of H100 time per fold against a single encoder pass plus C cosine comparisons. Its advantage decays with label-space granularity until, on a 48-intent taxonomy, it falls below the fine-tuned baseline it was meant to displace. Neither is a frozen multilingual encoder a substitute: it matches us on generic review sentiment and falls 6.8× behind on legal text.

04

It names a methodological blind spot — and then removes it as an excuse

Work in this area routinely omits m, the size of the set on which the threshold is selected, yet m governs the bound: Brands sits at m = 46 against a feasibility requirement of m ≥ 46, so its zero coverage is an artefact of sample size at the exact boundary. But we then ran the sweep that disentangles m from score quality, subsampling the calibration split from 5% to 100%: seven of eight corpora sit at exactly zero coverage at every fraction, including one at m = 21,016. Where the score does not separate, no amount of calibration data recovers coverage.

05

Breadth, in two languages, with the failures reported

Twelve Portuguese corpora and eight English ones under an identical protocol, five task families, three budgets, five draws each — not a single toy benchmark. Corpora were not filtered for favourable results: Brands and Eniac certify zero coverage for every method tested and are reported alongside the rest, as is Reuters, the one corpus across all twenty where the baseline certifies and we do not.

We did not invent SGR or prototypical fine-tuning. The contribution is the evidence — at scale, in two languages, with the two factors separated — that this is the combination that holds under extreme data scarcity (5-shot), and that the confidence function the field has been optimising is not what carries the effect.

Scope, as everywhere on this page: the detailed grid covers four corpora, though the ladder it summarises is measured on all twelve; its CE-encoder prototypes were built from validation examples rather than the training support, because the original checkpoints were not kept; we retrained the baseline saving the support features and the grid ordering does not change — over the CE representation, cosine remains worse than MSP on all four corpora. The ladder is monotone by median over the twelve Portuguese corpora and by mean over the eight English ones, not corpus by corpus.

Method

A pipeline of decoupled modules

Representation learning, scoring and risk-calibrated thresholding are independent stages, so the final rule applies regardless of what happens upstream — and any gain in confidence ranking translates directly into higher coverage at a fixed target risk.

DECOUPLED MODULE PIPELINE input text Encoder BERTimbau, 6 layers frozen tokens Episodic prototypes the operative ingredient latent space Geometric score κ(x) = maxc τ cos(z, μc) confidence SGR threshold binomial inversion at r* κ ≥ θ* κ < θ* Predict Abstain
Figure 1 — the pipeline. A Transformer encoder, an episodically trained prototypical space, a geometric confidence score, and an SGR thresholding step that selects θ* by inverting a binomial tail bound at target risk r* and confidence δ. Calibration never sees OOD data, so abstention on unseen classes is an emergent behaviour of the threshold, not a guaranteed one — which is what the open-world gap below measures.

1  Episodic prototypical representation

Class prototypes are support means over the label-adapted representation, μc = (1/K) ∑k z(xc,k), trained episodically with a prototype-repulsion term added to cross-entropy. The backbone we benchmarked (LAQDA) adds a label adapter and a query-neighbour augmentation step, both of which our ablation finds dispensable; the configuration we recommend is the plain prototypical encoder.

(1) z(x) = 0.1 · Aφ(Ĥ(x))CLS + 0.9 · Ĥ(x)CLS

Raising the adapter's share of that mixture ninefold, to 0.9, moves AURC by 0.003 — an order of magnitude inside the cross-fold standard deviation — so the ablation result is not an artefact of the inherited weight.

2  Geometric confidence

Confidence and prediction are read off the same geometry, with an independently set temperature τ. Because τ > 0 is a strictly monotone rescaling, it provably cannot change AURC or the accepted set at any threshold. It is, as the attribution grid shows, not the mechanism either: over a prototypically structured space, swapping this score for MSP moves AURC by 0.001. Treat it as the natural read-out of the geometry, not as the thing that produces the gain — on an unadapted representation MSP is in fact the better read-out.

(2) κ(x) = maxc τ cos(z(x), μc)     f(x) = argmaxc cos(z(x), μc)

3  The adopted thresholding rule (SGR)

Given a threshold-selection set Sm of size m, a target risk r* and a confidence level δ, SGR binary-searches over quantiles of κ, inverting a binomial tail at each candidate and keeping the most permissive threshold that still meets the target:

(3) b* = sup{ b : ∑j=0kθ C(mθ, j) bj(1−b)mθ−j ≥ δ′ }
(4) θ* = min{ θ : b*(mθ, r̂θ, δ′) ≤ r* },   δ′ = δ / ⌈log2 m⌉

The constraint is distribution-free and holds for any score — but validity and usefulness come apart. The bound stays valid under a poor score while the coverage it yields collapses to zero. That gap is the paper's subject.

Why the certificate is about ID risk only

For an OOD instance every acceptance is by definition an error, and calibration never sees OOD data. A threshold tuned for ID errors therefore rejects unseen classes only insofar as the geometry places them at lower confidence. We measure that behaviour; we do not certify it — and the measurement is unflattering.

Results

Selective classification across twelve corpora

All post-hoc baselines share an identical fine-tuned encoder and differ only in the confidence score. The comparison below is against MSP, the strongest of the seven, at its oracle-swept configuration — selected on the test set, an advantage we do not grant our own pipeline.

Area under the risk–coverage curve, 5-shot

Mean over 5 folds. Lower is better. Computed over the full ID + OOD mixture.
Metric
Ours (prototypical + SGR) MSP, oracle-swept (best post-hoc baseline)
Table 1. Selective classification at 5-shot, mean over 5 folds. E-AURC subtracts the AURC an optimal ranker would achieve at the same error rate. Ours is better on all three metrics on all twelve corpora; E-AURC discounts accuracy only partly, and the accuracy-invariant comparison is in the attribution grid.
MSP (swept)Ours
Corpus%OODAURCE-AURCAcc (ID)AURCE-AURCAcc (ID)

Accuracy is over ID inputs only; AURC and E-AURC are over the full ID + OOD mixture, in which every OOD input counts as an error.

The largest gains appear where conventional few-shot classifiers fail most. AURC drops 92% on HateBR and 92% on B2WReviews, 89% on RePro and 88% on Brands. The single largest gain is legal: RulingBR goes from 0.450 to 0.073 (−84%), with the study's largest ID accuracy gain, +45 points. The advantage is not confined to the two tabulated risk levels — it holds across the whole risk–coverage curve — and it is not explained by surface form, since a TF-IDF classifier on the same support is weaker than either method.

It is also the corpus with the widest spread across folds, and the mean alone describes it poorly. On RulingBR our selector's AURC is 0.073 ± 0.069 with a median of 0.045: one fold records 0.201 against 0.012 for the best. Anyone reproducing this should expect that variation. On the other ten corpora the coefficient of variation sits between 0.11 and 0.32 and the mean describes the distribution well; Brands is the other skewed case, but because two folds have AURC exactly zero, not because of a bad tail. The AURC spread does not carry over to certified coverage: on RulingBR it is 0.943, 0.997, 0.956, 0.875 and 0.943 fold by fold.

That dispersion belongs to the data, not the optimiser — and optimisation variance, once measured, is not negligible. We retrained with a second seed on all twelve corpora and a third on three of them, holding the support set fixed, so that only initialisation and episodic sampling vary. On IntentPT and RulingBR the ordering of the five folds by AURC is identical across the three seeds: the hard fold stays hard under a different initialisation. At the median across the twelve, the shift between seeds is 0.0015 AURC, or 9.7% of the between-fold standard deviation — but the range runs from 0.3% to 100.5% on Buscape, where optimisation noise equals the between-fold spread. Recognasumm is the second such case, at 0.0142.

Two consequences we would rather record than omit. The shift between seeds exceeds the effect of switching the score on four of the twelve corpora, so the 0.0023 denominator of the representation/score ratio should be read as an upper bound on the design's resolution, not as a measurement of the effect. And the numerator's margin is not uniform: the ratio between the representation effect and the seed shift has a median of 51 but falls to 3.5 on Buscape. An earlier version of this analysis, on three corpora only, reported a minimum of 19 — extending it to twelve gave a worse result, and it is the twelve-corpus one we report.

All seven post-hoc baselines

AURC across all seven post-hoc baselines

Twelve corpora × eight methods, 5-shot. Darker means worse.
Table 2. AURC ↓ at 5-shot, mean over 5 folds, all seven post-hoc baselines at their published configuration. Bold marks the best baseline per row; ours is lower than every baseline on every row. †RePro averages four folds: fold 01's archived report records its accuracy inverted relative to its own saved logits, so we exclude it uniformly across all seven scorers.
CorpusMSPEnergyMahal.kNNGradNormReActConjNormOurs

Attainable coverage at a target risk

Fraction of inputs the system answers rather than defers, with the threshold selected in-sample. These are operating points, not certificates — see Certification.
Budget
Target risk
OursMSP
Table 3. Attainable coverage at target risk r*, mean over 5 folds, δ = 0.05, threshold selected on the evaluation split by the same procedure for both methods. MSP is oracle-swept at 5-shot and at its published configuration at 10-shot, since the sweep covers 5-shot only. “≈0” is below 0.01. Each cell reads MSP → Ours.
5-shot10-shot
CorpusSGR@5%SGR@10%SGR@5%SGR@10%

Under that rule ours leads MSP on ten of twelve corpora at both risk levels, with mean coverage 0.50 vs. 0.13 at 5% and 0.62 vs. 0.24 at 10% (p = 0.00195 by a Wilcoxon signed-rank test at both levels over n = 10 pairs, after dropping the two corpora where both methods cover ≈0; with ten pairs, 2/210 = 0.00195 is the smallest attainable value. On AURC, where no pair ties, the same test gives 0.00049 with n = 12).

Two rows that measure the protocol, not the method — and the sweep that proves it

Brands collapses to ≈0 coverage despite 0.978 accuracy because its threshold-selection set sits exactly at the feasibility boundary (m = 46 against a bound of m ≥ 46 at 10% risk). Eniac collapses because its held-out OOD mass floors the empirical risk over the mixture at ≈20%.

That invites an obvious objection — that the baselines' collapse is really a small-m artefact too. We ran the sweep that settles it, subsampling the calibration split to 5, 10, 25, 50 and 100% of its size at fixed representation and score: seven of eight corpora sit at exactly zero coverage at every fraction, including Recognasumm at m = 21,016 and IntentPT at m = 2,514, both orders of magnitude above the feasibility bound. Only B2WReviews, which is not degenerate to begin with, rises with m (0.25 → 0.51). Where the score does not separate correct from incorrect, no amount of calibration data recovers coverage.

Attribution

Representation or score?

The headline comparison changes two things at once: the representation and the confidence score. A 4×2 grid separates them, evaluating two scores over four representations — a frozen multilingual encoder, a conventionally fine-tuned one, SetFit, and ours. The answer is not the one we expected.

MSP score
Cosine score
CE encoder
0.450Table 1 baseline
0.518worse than MSP
Prototypical
0.072 
0.073our pipeline
Across a row — change only the score: Δ 0.001 Down a column — change only the representation: Δ 0.378
Figure 2 — two rungs of the ladder on RulingBR (AURC ↓, 5-shot, mean over 5 folds). The two prototypical cells come from the same forward pass, so they share a representation and a prediction, and differ only in how confidence is read.

The 4×2 grid, corpus by corpus

Four representations × two scores. The bars move down the rungs, not across each pair.
Metric
Table 4. The full 4×2 grid, 5-shot, mean over 5 folds. Frozen is the better of BGE-M3 and multilingual-E5 per corpus, chosen on test, and CE is at its best of three learning-rate settings, also chosen on test — oracle advantages granted to the baselines, never to us. SetFit is at its best of three training settings on B2W and TuPy, chosen by validation AURC. Changing the representation moves the numbers; changing the score within a representation barely does.
Repr.ScoreB2WIntentRulingTuPy
AURC ↓
FrozenMSP0.0140.2170.4980.229
Frozencos0.0170.2600.5660.322
CEMSP0.0480.2830.4500.167
CEcos0.0640.2930.5180.343
SetFitMSP0.0400.2720.3470.294
SetFitcos0.0280.2680.3540.413
protoMSP0.0040.1150.0720.068
protocos0.0040.1170.0730.066
Correct-vs-incorrect AUROC ↑ (over ID inputs, accuracy-invariant)
FrozenMSP0.8480.8290.6880.651
Frozencos0.7870.7390.5610.512
CEMSP0.8110.8060.7040.600
CEcos0.6910.7760.6190.598
SetFitMSP0.8190.8260.7510.620
SetFitcos0.8440.8280.7370.538
protoMSP0.8980.8760.9360.781
protocos0.9170.8630.9290.779

The representation carries essentially the entire effect and the score carries almost none. Holding the score fixed at MSP, moving from the CE baseline to the prototypical encoder takes RulingBR from 0.450 to 0.072; holding the representation fixed, switching from MSP to our cosine score moves it from 0.072 to 0.073. Across twelve corpora the mean absolute difference between the two scores over the prototypical representation is 0.0036 (median 0.0023), with no consistent winner. The accuracy-invariant panel says the same: over the prototypical representation the two scores differ by at most 0.019.

Where the two representation-effect numbers on this page come from. Holding the score fixed at MSP and changing only the geometry — from the CE encoder to the prototypical one, over the same backbone — the median across the twelve corpora is 0.056 in E-AURC and 0.091 in correct-vs-incorrect AUROC. Against 0.0023 and 0.014 on the score side, those give the ratios of 24 and 6.5 that bound the range. The frozen-encoder contrast does not serve here, because it also swaps the base architecture. One identity is worth stating, because we rely on it: since cos and MSP share the argmax, the error rate is the same under both read-outs and the term E-AURC subtracts cancels in the difference, so Δ E-AURC = Δ AURC exactly for the score contrast — which is why 0.0036 and 0.0023, measured in AURC, also hold in E-AURC. The identity does not hold for the representation contrast, where accuracy changes, so the E-AURC column on that side has to be computed.

It is worth bounding what that cell establishes. The two scores contrasted there are functions of the same similarity vector over the same prototypes, so a small gap is partly expected by construction. The score-variant appendix widens the set to five estimators, including X-Mahalanobis, which reads the same geometry through cross-layer covariance and is the only one to stand apart — and even it sits 0.0096 from cosine, an order of magnitude below the representation effect. All seven scores of the full grid have now been evaluated over the prototypical representation, and the result separates three claims.

Energy and ReAct converge; the four structurally distinct ones do not. The Energy estimator demonstrates a median difference of +0.0003 relative to cosine across all twelve corpora, while ReAct exhibits a +0.0001 margin, resulting in a spread of 0.0033 among the four estimators that leverage prototypical geometry — comparable in magnitude to the 0.0023 margin of the cosine/MSP pairing. Conversely, the four structurally distinct scores remain behind: kNN at +0.0439 from cosine at the median, ConjNorm at +0.0101, Mahalanobis at +0.0097 and GradNorm at +0.0069, each worse than cosine on eleven or twelve of the twelve corpora. Given that all preserve the prototypical argmax output, these discrepancies constitute E-AURC differences, directly comparable to the 0.056 representation effect. When considering the spread across all eight estimators, the impact ratio between representation and score is reduced from 24 to approximately 1.2.

Adapting the representation halves the price of choosing badly. The same seven post-hoc scores show a median spread of 0.1060 AURC over the cross-entropy encoder and 0.0460 over the prototypical representation — a 2.3-fold reduction. This is the attribution-ladder claim verified outside the family of scores that motivated it, rather than only among readings of the same similarity vector.

Two observations bound this counterexample without invalidating it. The deficit is neither uniform nor a property of the latent representation: Mahalanobis and ConjNorm actually beat cosine on RulingBR (0.071 against 0.073) and fall well behind on Buscape (0.136 and 0.121 against 0.057). What the four share is dependence on an input the regime rations — a dense bank for kNN, an estimable covariance for Mahalanobis, a sampled partition constant for ConjNorm — and all must estimate it from CK support examples. Consequently, the empirical evidence supports the narrow claim: over a task-adapted geometry, estimators provided with sufficient operational conditions converge with one another and the cost of a bad choice halves — without the choice becoming indifferent.

The ladder: how much the score matters, by how much the space was adapted

Score sensitivity falls with adaptation

Absolute gap between the two scores read off the same representation, in correct-vs-incorrect AUROC over ID inputs. Lower means the choice of score matters less.
Portuguese (median, 12 corpora) English (mean, 8 corpora)
Table 18. Absolute gap between the two scores read off the same representation, in correct-vs-incorrect AUROC over ID inputs. Portuguese is the median over the twelve corpora, English the mean over the eight.
RepresentationPortugueseEnglish

From no adaptation to episodic prototypical training the gap falls from 0.107 to 0.014, a factor of about eight, and the English replication gives 0.110 to 0.014 over its eight corpora. Which scalar is read off the space matters most exactly where the space was never shaped for the task, and becomes almost irrelevant once it was.

Two qualifications we state rather than round away. The step from a frozen encoder to a conventionally fine-tuned one is not resolvable at this sample size (0.107 against 0.099); only the two steps involving representation learning explicitly structured for the task — contrastive and episodic — survive, and those are separated by a factor of four. And the ladder is monotone by median over the twelve Portuguese corpora and by mean over the eight English ones, not corpus by corpus.

Decomposing the gap by score says why, and the answer is one-directional. This decomposition changes scope, and we say so: it is the mean over the four deep-dive corpora of the 4×2 grid, not over the twelve behind the medians above. Across the four rungs MSP rises 0.754 → 0.730 → 0.754 → 0.873 while the cosine score rises 0.650 → 0.671 → 0.737 → 0.872. MSP is at least as good as the cosine score on every representation in both languages; the gap closes because adaptation rescues the geometric read-out until it merely ties. The honest statement is not that our confidence function is better — it never is — but that on an adapted representation the choice of read-out stops costing anything.

REFINEMENT 01

A frozen encoder: split verdict

On B2WReviews it attains the lowest AURC of the four representations (0.014) and certifies 0.98 coverage without a single gradient step. But on RulingBR it collapses to 0.498 — worse than a conventionally fine-tuned BERTimbau and 6.8× behind ours — certifying no coverage at all. Pretrained multilingual geometry suffices for generic high-resource tasks and is no substitute on specialist ones.

REFINEMENT 02

SetFit: membership in the family is not sufficient

Contrastive fine-tuning recovers part of the effect where the baseline is weakest (RulingBR 0.450 → 0.347) but not elsewhere, and even oracle-tuned stays 7× behind us on B2WReviews and 4.5× on TuPy, certifying nothing on TuPy where we certify 0.23. How the representation is restructured matters as much as whether it is.

REFINEMENT 03

At one shot, SetFit collapses entirely

At K = 1 it is worse than the CE baseline on three of four corpora and certifies zero coverage everywhere, because its positive pairs degenerate into self-pairs when a class has a single example — where an episodic prototype is still well defined, if noisy. We state the mechanism as a hypothesis: we swept hyperparameters only at 5-shot.

This does not collapse into “fine-tuning yields a better classifier”

Correct-vs-incorrect AUROC is computed over ID inputs only and is invariant to how many the model gets right. Against the Table 1 baseline it rises from 0.704 to 0.929 on RulingBR, 0.600 to 0.779 on TuPy, 0.811 to 0.917 on B2WReviews and 0.806 to 0.863 on IntentPT. Matching accuracy by subsampling — discarding correct predictions from the more accurate model until the two match, against the same swept baseline as the main tables — gives AURC 0.029 vs. 0.048 on B2WReviews, 0.252 vs. 0.283 on IntentPT, 0.308 vs. 0.450 on RulingBR and 0.102 vs. 0.167 on TuPy — better on all four. Restricted to ID inputs: 0.029 vs. 0.048, 0.157 vs. 0.182, 0.206 vs. 0.340 and 0.102 vs. 0.167. The random discard is repeated 50 times per fold. The ranking advantage survives when the accuracy advantage is removed entirely, on either population.

But the improvement is a property of the learned geometry, and attributing it to our choice of confidence function — as an earlier framing of this work did — is not supported by the data.

Certification

From attainable to certified

On all twelve corpora we re-ran the pipeline selecting the threshold on the ID validation split and applying it unchanged to test. That split is disjoint from the support set and the evaluation set — but not from the procedure that picked the checkpoint, since early stopping monitors the loss on that same split. The Geifman & El-Yaniv guarantee requires independence from the classifier, so we redid the certification on a split no model-selection step ever touched: half the test set calibrates, the other half measures. Mean coverage falls from 0.569 to 0.538 at 5% risk, and realised risk exceeds the target in 26 of 1,908 cells — a 1.4% violation rate against the nominal δ = 5%. That is the number backing the word “certified”; the table below keeps the original calibration, which is identical for both methods and so favours neither.

Certified ID coverage, twelve corpora

Threshold fitted on the ID validation split and applied unchanged to test.
Target risk
Table 5. Certified ID coverage and realised ID selective risk at target risk r*, 5-shot, mean over 5 folds, δ = 0.05, all twelve corpora. MSP is at its swept configuration. “—” marks an infeasible target, where no threshold satisfies the bound and no risk is defined.
r* = 5%r* = 10%
CorpusMSP covriskOurs covriskMSP covriskOurs covrisk

Of the 120 corpus-fold-level cells, 97 admit a feasible threshold and the certificate is violated in two of those 97 — a rate of 2.1% against the nominal δ = 5%. Both violations are third-decimal (Buscape fold 02 realises 0.0503 at a 5% target; HateBR fold 01 realises 0.1012 at 10%). Infeasible cells are excluded from the denominator and the 97 cells are not independent, so this is a consistency check against δ, not a binomial test of it. A low mean can mean either small coverage on every fold or reasonable coverage on some and none on the rest: at 5% risk our selector admits a feasible threshold on 5 of 5 folds in eight corpora, 3 of 5 on Buscape and TuPy, and none on Brands and Eniac, while MSP is at 0 of 5 on six corpora.

The ordering survives certification and widens. Our selector certifies more coverage on ten of twelve corpora at 5% and eleven of twelve at 10%, with mean certified coverage 0.569 against 0.136 and 0.731 against 0.278. MSP admits no feasible threshold at 5% on six of the twelve — Brands, Eniac, HateBR, IntentPT, RulingBR and TuPy — while ours certifies non-zero coverage on four of those six. The two exceptions are the two corpora where every method collapses, and we report them rather than omitting them: a certificate that returns zero coverage is the procedure behaving correctly. The two rows nonetheless reach zero for different reasons. On Eniac the score cannot support the target, with m = 580, comfortably above the floor. On Brands the calibration set has m = 45, below the binomial floor of m ≥ 97 at 5% risk: coverage there is zero for want of sample, and a perfect classifier would be zero too. Resampling the observed score distribution up to m = 500, the same model would certify 0.95.

Infeasibility is a property of the confidence values, not of the search: scanning every threshold in the calibration split, the tightest bound MSP reaches is b* = 0.211 on RulingBR and 0.068 on IntentPT. The correction used in that scan should be stated: it keeps δ′ = δ/⌈log2 m⌉, sized for the candidates a binary search visits, rather than δ/m, which is the correction proper to a scan over all m quantiles — the more permissive of the two, so the reported b* is an optimistic floor. Taken to the limit, with no correction at all the floor drops to 0.149 on RulingBR and 0.049 on IntentPT. RulingBR's infeasibility at 5% is therefore independent of the correction; IntentPT's is not — there it depends on the multiplicity of the search being counted, as any valid certificate must.

Against SetFit, on the deep-dive corpora

Certified coverage against SetFit

Four deep-dive corpora, three methods. Missing bars are infeasible targets.
Target risk
Table 6. Certified coverage with a SetFit row, 5-shot, mean over 5 folds. SetFit is at the library default, at the score validation AURC selects per corpus; its own configuration sweep, also selected on validation, reaches 0.80 on B2W. An infeasible target — no threshold meets the risk target — is reported as 0.00 coverage, the same convention as Table 5.
r*MethodB2WIntentRulingTuPy
5%CE + MSP0.590.000.000.00
SetFit0.400.070.070.00
Ours1.000.630.940.23
10%CE + MSP0.820.240.000.00
SetFit0.770.210.190.00
Ours1.000.861.000.65

The choice of risk-control rule is not neutral either

SGR vs. CRC vs. Learn-then-Test

Certified ID coverage over our own representation.
Target risk
Table 7. Certified ID coverage under three risk-control rules over our representation, 5-shot, mean over 5 folds, calibrated on the ID validation split. CRC is never below SGR and is strictly above it on nine of twelve at both levels; LTT is below SGR on eight.
r* = 5%r* = 10%
CorpusSGRCRCLTTSGRCRCLTT

CRC or SGR: the choice is conditional, not a general recommendation

CRC obtains its advantage by not paying for the threshold search, which is valid only when selective risk is monotone in the threshold — a condition selective risk does not guarantee in general. That condition is testable on the calibration splits we already hold, so we measured it rather than assuming it: over 200 threshold quantiles per fold, the mean largest violation across the twelve corpora is 0.0061, with seven below 0.003 and a worst case of 0.0207 on Eniac. At these magnitudes the empirical risk is monotone up to sampling noise.

The gap is concentrated where SGR's union correction costs most: on Brands, whose m = 45 sits one point below SGR's feasibility bound at 10%, CRC turns a certified 0.000 into 0.933 — so some of what we attribute to small m is attributable to the correction instead. LTT, which controls the search with a Bonferroni correction, falls below SGR on eight of the twelve corpora at 5% risk, ties on three, and beats it only on Buscape.

The caveat that decides between the two rules is not monotonicity but the form of the guarantee. SGR and LTT give PAC guarantees — they bound Pr(R > r*) by δ over the draw of the calibration set — whereas CRC controls risk in expectation over that same draw. Demanding expectation is strictly less than demanding high probability, and part of CRC's extra coverage follows from demanding less, not from being a better procedure. The practical rule is a single one: use CRC when the commitment is to the average risk of the system across many deployments and m is small enough that SGR's union correction collapses coverage; use SGR when the commitment is to this deployment, on this calibration set. CRC covers more because it promises less, and does not replace SGR where the stronger promise is the one required.

The number a practitioner most needs: certified ≠ safe under shift

Applied to the full ID + OOD mixture, the same certified thresholds accept 0.86 of RulingBR at r* = 5% — and because accepted OOD inputs are errors by definition, the realised risk becomes 15.2% against a 5% target, and 23.2% against 10%. On IntentPT the pair is 10.0% and 20.7%. Equivalently, the 5% threshold admits 52% of all OOD inputs on RulingBR.

The mechanism is arithmetic, and it runs opposite to intuition: the threshold is chosen to meet an ID risk target, so a cleaner ID geometry earns a more permissive threshold, and a permissive threshold admits OOD regardless of how well the score ranks it. RulingBR has the study's best ID/OOD separation (AUROC 0.893) yet admits 52% of OOD; IntentPT separates worse (0.797) and admits 19%. Good ID representation buys coverage, not open-world safety, and the two can move in opposite directions on the same corpus.

Cross-lingual replication

Eight English corpora, identical protocol

Every headline result above is Brazilian Portuguese, so the attribution could be a property of the language, the corpora or the encoder. We repeated the pipeline and the grid on eight English corpora with bert-base-uncased substituted for BERTimbau throughout — same folds, same support, same calibration split, same SGR code, so that the language changes and nothing else does.

English corpora: CE baseline vs. ours

Identical protocol, bert-base-uncased in place of BERTimbau.
Metric
Table 8. English corpora, 5-shot, mean over 5 folds. The CE baseline receives the same three-configuration oracle learning-rate sweep, chosen on test, that it receives in Portuguese. The last two columns are ID/OOD separation AUROC (correct-vs-incorrect AUROC is the ladder in Table 18); it is undefined for SST-2, whose two ID classes leave no OOD partition.
AURC ↓AUROC ID/OOD ↑
CorpusCEOursΔ%CEOurs

Median reduction 56.8% against 66.9% in Portuguese; Wilcoxon p = 0.0078, the smallest value attainable at n = 8, with a median paired difference of 0.127 (95% bootstrap [0.077, 0.191]). The smaller margin is the expected direction: bert-base-uncased is a stronger starting point for English than BERTimbau is for the specialist Portuguese domains.

Certified coverage in English, r* = 5%

Infeasible folds contribute zero coverage to the mean.
Table 9. Certified ID coverage at r* = 5% on the English corpora, 5-shot. “feas.” counts the folds admitting a feasible threshold; infeasible folds contribute zero to the mean. The 10% level was not stored by these runs.
MSPOurs
Corpuscovfeas.covfeas.

The one case in either language where the baseline certifies and we do not

On Reuters the ordering reverses: MSP finds a feasible threshold on one of five folds and our selector on none. StackOverflow is infeasible for both. These are the two English corpora with the least favourable ratio of calibration size to label cardinality (34 and 251 ID classes), and the two where our AURC advantage is smallest (29.8% and 21.6% against 53–91% elsewhere) — consistent with the certificate being hardest exactly where the score separates least, though with two corpora we cannot distinguish that from chance. We report it rather than restricting the table to the seven favourable rows.

Analysis

Budgets, OOD, ablation, backbones, LLMs

Effect of the shot budget

AURC versus budget K

Mean over 5 folds, bars are one standard deviation. MSP is at its published configuration at all three budgets, since the oracle sweep covers 5-shot only. Note the per-panel vertical scale.
OursMSP
Table 17. AURC ↓ versus budget K across the twelve corpora, mean ± std over 5 folds. MSP is at its published configuration at all three budgets. Every cell is recomputed from the saved test logits, which removes the fold exclusion described in Table 2: this is why the RePro cell at 5-shot is 0.083 over five folds rather than 0.070 over four.
Corpus / method1-shot5-shot10-shot

Across the 36 cells of the grid our method has a lower AURC than MSP, with a median reduction of 73.1% at 1-shot, 78.2% at 5-shot and 59.8% at 10-shot. The advantage is largest under extreme scarcity, consistent with the mechanism: at K = 1 the unadapted prototype is the single support embedding, with no averaging to damp representation noise. The widest absolute margin is RulingBR at 1- and 5-shot (0.68 and 0.41 AURC) and Eniac at 10-shot (0.18). The reading that the margin shrinks with the budget holds for the absolute difference; in relative terms the ordering is not monotone.

OOD detection under class holdout

OOD detection under class holdout

The five corpora with held-out classes, 5-shot.
Metric
Table 10. OOD detection at 5-shot on the five corpora with held-out classes, means over 5 folds.
AUROC ↑FPR@95 ↓
CorpusMSPOursMSPOurs
Eniac0.5530.6850.9280.803
Recognasumm0.6570.7700.8640.794
IntentPT0.6790.7970.8710.771
MMLU-PT-BR0.7020.8040.8610.787
RulingBR0.5550.8930.9310.472

Rejecting near-OOD from few-shot prototypes is only partial — FPR@95 stays above 0.77 on four of the five corpora, and only on RulingBR does it fall below one half (0.472). But that difficulty is specific to the class-holdout protocol, which produces near-OOD by construction: held-out and retained classes share domain, register and much of the vocabulary. Substituting text from an unrelated domain for the held-out classes, with the checkpoint untouched, raises AUROC to 0.986 on Eniac and 0.997 on RulingBR. Far-OOD is close to saturated for this geometry.

Ablation: which ingredient does the work?

Ablation on IntentPT

Whiskers are one cross-fold standard deviation.
Table 11. Ablation on IntentPT (48 classes), AURC ↓, mean ± std over 5 folds. All four configurations beat every post-hoc baseline on this corpus by a wide margin (0.105–0.119 against 0.283 for the swept MSP).
Configuration1-shot5-shot10-shot
Full (label adapter + QDA).138 ±.020.117 ±.027.113 ±.026
w/o QDA sampler.134 ±.017.105 ±.023.103 ±.021
w/o label encoder.139 ±.015.119 ±.026.113 ±.024
Plain prototypical.131 ±.018.107 ±.024.106 ±.024

The label adapter is dispensable — and that makes the finding stronger

A plain prototypical encoder — no label semantics, no query augmentation — reproduces full LAQDA on all twelve corpora, mean AURC 0.075 against 0.074, differing by 0.004 on average and lower on six of twelve. The operative ingredient is episodic prototypical fine-tuning — and, per the attribution grid, not the score it is read with either. So the coverage gains do not depend on a particular adapter, and the lever is available to any practitioner already using prototypical networks.

How much depends on the encoder?

Backbone and read-out controls

Same folds, support sets and calibration protocol.
Table 12. AURC ↓ at 5-shot, mean over 5 folds: the swept MSP baseline, our pipeline on BERTimbau, the same pipeline on XLM-R base, and a logistic probe on frozen embeddings at the better of BGE-M3 and multilingual-E5 per corpus chosen on test. All four share folds, support sets and calibration protocol.
CorpusMSPOursXLM-RFrozen + probe

Two controls, and both cut against an unqualified reading. The mechanism is not encoder-independent: XLM-R still beats the swept baseline on nine of twelve corpora but its mean AURC is 0.194 against BERTimbau's 0.074, with the losses concentrated on colloquial Brazilian review and social-media text (TuPy 0.550 vs. 0.068, Buscape 0.308 vs. 0.047) — precisely the register a Portuguese-specific encoder is best placed to serve. On the three corpora where the encoders come close, the difference sits within the spread across folds (Recognasumm 0.137 vs. 0.136, RulingBR 0.075 vs. 0.073, IntentPT 0.114 vs. 0.117). And a trained read-out over a frozen space buys roughly what supervised fine-tuning of the whole encoder does (mean 0.203 against 0.219) but does not close the gap to episodic training, failing on exactly the specialist corpora where the centroid read-out fails.

Against a prompted 7B LLM — and against TF-IDF

Prompted LLM and lexical baseline

Three corpora spanning 2, 10 and 48 classes.
Metric
Table 13. A lexical baseline (TF-IDF + logistic regression on the same K examples) and prompted Qwen2.5-7B-Instruct against our pipeline and the swept MSP baseline, 5-shot, means over 5 folds. The LLM reads a categorical distribution from answer-token logits under the same folds, support sets and calibration split.
CorpusMethodAURC ↓E-AURC ↓Acc ↑
B2WReviews
2 classes
TF-IDF0.1050.0790.781
MSP0.0480.0360.855
Qwen2.5-7B0.0110.0110.973
Ours0.0040.0030.971
RulingBR
10 classes
TF-IDF0.6480.3100.388
MSP0.4500.1950.495
Qwen2.5-7B0.1920.0780.735
Ours0.0730.0220.944
IntentPT
48 classes
TF-IDF0.4150.1450.457
Qwen2.5-7B0.3850.1540.512
MSP0.2830.1170.607
Ours0.1170.0570.849

Three things follow. The premise we were prepared to assume is false: answer-token logits do give a usable selective signal, and on B2WReviews the LLM reaches attainable coverage 1.00 at 5% risk against MSP's 0.59. That corpus isolates the central claim more cleanly than any of our own ablations — the LLM is more accurate there and still worse at AURC, so our entire advantage is in error ranking. Second, that competitiveness decays with label granularity until, on a 48-intent taxonomy, it falls below the fine-tuned baseline; both gaps survive E-AURC. Third, the cost asymmetry does not narrow: 2,756 ± 537 seconds of H100 time per RulingBR fold, because every input needs attention over the whole support set in context.

A fourth corpus separates two distinct failures. On MMLU-PT-BR (45 classes) the LLM's ID accuracy collapses to 0.089 — yet its E-AURC, which discounts accuracy, is 0.106, essentially MSP's 0.094. The confidence is still rankable; what the model cannot do is recover a labelling convention (subjects separated by education level) that our pipeline infers from the same five examples per class.

Score variants and calibration

Confidence-score variants over the same representation

Averaged across the twelve corpora at 5-shot.
Table 14. Mean AURC and mean gain over MSP for each confidence-score variant over the same prototypical representation, averaged across the twelve corpora at 5-shot. The comparator is MSP in the oracle-swept configuration (mean 0.2181 across the twelve), the same as in Table 1. Four decimals because the spread across variants sits in the third: all five fit within 0.011 of AURC. No variant is universally superior, so the main tables report the base cosine score rather than the per-corpus best — which would amount to selecting on the test set.
ScoreAURC ↓Gain vs. MSP ↑
Margin0.0723+0.1458
TempScale0.0725+0.1456
MC Dropout0.0727+0.1454
Cosine (reported)0.0737+0.1444
X-Mahalanobis0.0833+0.1348

The narrow spread is the attribution seen from another angle: once the space is structured, the choice of score is close to immaterial. The gains are in ranking, not calibration — the cosine score is not a probability and its ECE is poor (0.316 on B2WReviews against 0.038 for MSP in its published configuration — the one comparison on this page that does not use the swept MSP, because calibration and ranking do not move together under the sweep). Temperature scaling of the same similarities recovers calibration and even improves on MSP there (0.010), though not uniformly (0.115 on IntentPT against 0.046). Prefer TempScale when calibrated probabilities are needed downstream.

Calibration set size m

Calibration set size m, log scale

Dashed lines mark the feasibility bounds: m ≥ 46 at 10% risk, m ≥ 97 at 5%.
Table 15. Threshold-selection set size, fold 01. The twelve-corpus comparison runs select the threshold on the ID test split; the certified runs use the ID validation split. m spans three orders of magnitude. Feasibility with zero observed errors requires m ≥ 46 at r* = 10% and m ≥ 97 at 5%.
CorpusID valid.m (ID test)Certified cov @5%, subsampled 5% → 100%

The last column is the coverage-versus-m sweep at fixed representation and score, mean of five draws per fraction. If the collapse were a sample-size artefact each row would rise from left to right; only B2WReviews does. Dashes mark corpora not included in the sweep.

Scope — what these numbers do not establish

Only the certified runs carry the guarantee. Every attainable-coverage figure selects the threshold on the evaluation split and is an upper bound on what SGR could certify over that same population. The two are not comparable term by term and neither bounds the other, since they are measured over different populations.

The OOD protocol is class holdout, which produces near-OOD by construction and covers neither far-OOD nor temporal or adversarial drift. The main tables use a single seed (42), so the standard deviations there measure sensitivity to the data partition and not to initialisation; we measured what that omits by retraining with a second seed on all twelve corpora, and a third on three of them (see below). The detailed grid covers four corpora, though the ladder it summarises is measured on all twelve, and its CE-encoder prototypes were built from validation examples rather than the training support — a gap we closed by retraining the baseline with the support features saved, without changing the grid ordering. The prompted-LLM comparison covers one model size and one elicitation protocol; the LLM has very likely seen these public corpora in pretraining, a confound that favours it. Both base-sized encoders are BERT-family, and whether the gains persist with larger encoders remains open.

Reproducibility

Running the framework

Source code, the split configuration (ood_splits.json), the benchmark scripts and the per-fold results underlying every table are in the repository. Hyperparameters are fixed reference defaults, held constant across all twelve corpora and three budgets — the baselines, not our pipeline, receive the oracle sweeps.

1

Environment

uv for a native install (recommended on cluster / DGX), or Docker with GPU support.

make setup-env
source .venv/bin/activate
# or: make laqda-install && docker compose up laqda
2

Deterministic ID / OOD splits

Generates configs/ood_splits.json, mapping every corpus and fold. The ID/OOD split is resampled per fold, so the five folds capture the variance of which classes are held out.

python main.py
3

Train — prototypical encoder with SGR

Without --use_sgr the model runs at 100% coverage. With it, a critical threshold θ* is estimated and stored, and the model abstains below it.

python -m methods.laqda.cli.train \
    --dataset_dir data/datasets/datasets-br-nlp/intent/IntentPTCorpus/few_shot \
    --fold 01 --kshot 5 \
    --save_dir outputs/laqda_sgr/IntentPTCorpus/fold_01 \
    --use_sgr
4

Baselines under an identical encoder

One run benchmarks all seven post-hoc scores — MSP, Energy, Mahalanobis, kNN, GradNorm, ReAct, ConjNorm — on the same fine-tuned encoder.

python -m methods.baselines.cli.train_baseline \
    --dataset_dir data/datasets/datasets-br-nlp/intent/IntentPTCorpus/few_shot \
    --fold 01 --kshot 5 \
    --save_dir outputs/baseline/IntentPTCorpus/fold_01
5

Full benchmark and consolidation

SLURM scripts submit every algorithm across folds 01–05 and budgets K ∈ {1, 5, 10}; compare.py aggregates cross-fold means ± standard deviations and plots the reports.

bash scripts/run_all_pt.sh          # submit the grid
python scripts/compare.py --corpus RulingBRCorpus --plot
make test-baselines                 # scorer unit tests

Repository layout

configs/  model, method and OOD-split configuration
data/  datamodules and k-shot episodic samplers
methods/
  laqda/  label-aware encoder, QDA sampler, contrastive loss, train/infer CLI
  baselines/  MSP, Energy · distance/ Mahalanobis, kNN · sota/ GradNorm, ReAct, ConjNorm
  sgr/  binomial-inversion risk control and post-hoc thresholding
  metrics/  Acc, F1, ECE, AUROC, FPR@95, AUPR, AURC, E-AURC
scripts/  SLURM submission, aggregation, ablation and paper checks
tests/  mathematical and integrity tests for every scorer

Fixed configuration

Table 16. Reference defaults, deliberately not tuned per corpus.
ComponentSetting
EncoderBERTimbau base cased (pt) / bert-base-uncased (en), first 6 layers frozen
OptimisationAdamW, lr 2×10⁻⁵, weight decay 10⁻³, ≤100 epochs, early stopping (patience 7)
Episodes100 per epoch; K∈{1,5,10} support, ≤25 query per class
Adapterρ=15, τtr=20, τ=15, t=5 — all found dispensable
SGRδ=0.05, r*∈{0.05, 0.10}, binary search over calibration quantiles
Protocol5 folds; seed 42 in the main tables, with a second seed on all twelve corpora and a third on three of them to measure optimisation variance; ≈40% of classes held out as OOD on multiclass corpora (≈20% of test instances)
Citation

BibTeX

The entry below is a placeholder for the anonymised submission; it will be replaced by the published reference on acceptance.

@inproceedings{selective_risk_ptbr,
  title     = {Selective Classification in Few-Shot Text:
               The Gain Comes from the Representation, Not the Confidence Score},
  author    = {Anonymous},
  booktitle = {Under review},
  year      = {2026},
  note      = {Code: https://github.com/beatrizalmeidaf/selective-risk-framework}
}

Building blocks we adopt

SGR — Geifman & El-Yaniv, Selective Classification for Deep Neural Networks, NeurIPS 2017.
CRC / Learn-then-Test — Angelopoulos et al., conformal risk control.
LAQDA — Liu et al., 2024.  SetFit — Tunstall et al., 2022.  Prototypical Networks — Snell et al., NeurIPS 2017.
BERTimbau — Souza et al., BRACIS 2020.  XLM-R — Conneau et al., ACL 2020.
MSP — Hendrycks & Gimpel, ICLR 2017.  Energy — Liu et al., NeurIPS 2020.  kNN — Sun et al., ICML 2022.

Ethical considerations

Selective classification acts as a safety mechanism, but a stated target risk may induce automation bias: the guarantee is strictly conditioned on i.i.d. ID calibration data and offers no cover against distribution shift among accepted predictions — on our own certified runs, a 5% ID guarantee yields 15.2% realised error once unseen classes enter the stream. Abstention also raises fairness implications: if particular demographic or dialectal groups are under-represented in the support set, their inputs may be disproportionately rejected. We recommend auditing coverage stratified by subpopulation before any deployment. All datasets used are publicly available; the hate-speech corpora contain inherently offensive language, handled here exclusively for research aimed at improving moderation systems.