論文 / arXiv:2610.06851
Base Models Can Reason By Taking a Cue From Training Data
PLAIN SUMMARY / やさしい要約
この研究は、AIの返答の出だしにある短い言葉が、その後の「考え方」に影響しうるのかを調べています。特定の出だしを決め打ちすると、数学の成績が上がる例が示されました。さらに、学習用のデータを少し変えることで、どんな単語でも「考えるスイッチ」のように働かせたり、逆にその効果を消したりできると述べています。たとえるなら、会話の最初の一言が、その後の話し方の“モード”を切り替える合図になる、という見方です。安全性の面でも、出だしの違いで拒否するか従うかが変わりうる、という観察が紹介されています。
AIが専門用語を使わずに書いた解説です。内容の正確さは、下の原文の要旨で確認してください。
ABSTRACT / 要旨(原文)
In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while "Alright," raises Qwen3-14B's from 72% to 87%. Second, RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model. Third, we trace the reasoning effects of token cues to the training data. We perform causal data interventions to turn an arbitrary word, such as "chicken", into an effective reasoning cue, or remove an existing cue's effect. A similar edit makes the prompt instruction "Think duck duck goose" as effective as "Think step by step" at eliciting reasoning. We also find that the hidden state representations induced by different cues correlate with different document types from the training set. Finally, we extend our study of token cues with a case study in language model safety, finding that different cues elicit distinct refusal and compliance behaviors that correspond to different types of training data.
ここに表示しているのは原論文の要旨です。AIによる要約や解釈は含みません。
