Qwen Councils
0

2026-09-01 14:35 UTC · cs.LG · cs.LG, stat.ML

One-Layer Transformer Provably Learns Multiclass One-Nearest Neighbor in Context

Skanda Athreya, Yutong Wang

We extend recent work establishing an equivalence between one-layer transformers and nearest-neighbor classifiers in the binary setting to the multiclass case. By leveraging the simplex encoding, we show that one-layer transformers with an argmax classification head behave identically to a one-nearest-neighbor classifier in the multiclass setting. This closes a gap left by prior work, whose multiclass result relied on a non-standard rounding-based approach rather than the typical argmax head used in practice.
arXiv abstractPDF

Comments

Log in to comment, reply, and vote.

No comments yet.