SHIRITAS

論文 / arXiv:2610.06833

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

PLAIN SUMMARY / やさしい要約

同じ処理を何度も繰り返すタイプの言語モデルは、繰り返す回数ぶん計算が重くなります。この研究は、「最終的に落ち着く状態(固定点)に近づくなら、そこへ至る途中の細かい違いは重要ではない」という考え方を使い、計算を軽くする方法をまとめています。たとえば、目的地にほぼ着いているなら、最後の数歩だけ丁寧に追えば十分、というイメージです。学習や文章生成、さらに強化学習の更新でも、速くする余地があると述べています。

AIが専門用語を使わずに書いた解説です。内容の正確さは、下の原文の要旨で確認してください。

ABSTRACT / 要旨(原文)

Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; terminal key-value (KV) sharing for decoding with almost no loss in accuracy; a distilled student that prefills up to 1.79x faster; and RL updates that compute gradients from saved rollout states, 2x faster than backpropagating through the replayed trajectory. We therefore improve the two components of training that shape these fixed points: the depth prior and input injection. Fixed-depth training breaks KV sharing, and Huginn's broad depth prior supports sharing but dilutes supervision at the target depth more than sharing requires; we learn the prior from prediction feedback, with an entropy term that keeps it broad. Existing injection schemes let the state's component along the input amplify or cancel the injection; we remove this component with orthogonal injection. From 100M to 1.6B parameters, the learned prior and orthogonal injection lower perplexity at every scale relative to Huginn's prior and existing injection schemes, respectively. At 1.6B, the learned prior with a 3x smaller KV cache matches the downstream average of fixed-depth training with the full cache.

ここに表示しているのは原論文の要旨です。AIによる要約や解釈は含みません。

この論文を取り上げた記事