
Cognition が新たな SWE-2 モデルを発表。Fable 5.1 や GPT-Astra と競合する性能を誇ります。
## Japanese Translation: Cognition は、Fable 5.1 や GPT-5.6 Sol といったトップクラス競合に匹敵する最新コーディングモデルである SWE-2 を発表しました。SWE-2 は、大規模な Kimi K3 ベースモデル(パラメータ数 2.8 兆)での後学習により実現され、コストペナルティ付き報酬とファーストプリンシプルに基づくアプローチ、そして長さに基づく報酬ベースラインを採用してトレーニングを安定化させながら、多兆パラメータ領域への強化学習のスケーリングを達成しました。FrontierCode 1.1 Main では 50.0% のスコア(Fable 5.1 より僅か 1 ポイント下)を記録しながらコストは 64% 削減され、DeepSWE 1.1 では 73.0% を達成しました。単なるスコアを超え、SWE-2 は「エンジニアリング的判断」の優位性も示し、不要な迂回を避けることで初期コードエディットの中央値ステップ数を 48 から 18 に削減しました。また、モデルは安全性と信頼性を最優先しており、プロパガンダおよび検閲テストのうち 98% をパスしています。Devin Desktop と CLI 経由で即時利用可能(Web および Fusion では段階的展開中)の SWE-2 は、高パフォーマンス AI をアクセス可能な価格点で提供し、信頼性や安定性を損なうことなくソフトウェア開発サイクルの効率化を目的としています。 ## Text to translate: Cognition has introduced SWE-2, its most advanced coding model, which rivals top competitors like Fable 5.1 and GPT-5.6 Sol while offering significant efficiency gains. Achieved through post-training on the massive 2.8 trillion-parameter Kimi K3 base model, SWE-2 scales Reinforcement Learning to a multi-trillion parameter regime using cost-penalized rewards derived from first principles and a length-weighted reward baseline to stabilize training. On FrontierCode 1.1 Main, it scores 50.0% (just one point behind Fable 5.1) while being 64% cheaper; on DeepSWE 1.1, it achieves 73.0%. Beyond raw scores, SWE-2 demonstrates superior "engineering judgment," reducing the median steps to an initial code edit from 48 down to 18 by avoiding unnecessary detours. The model also prioritizes safety and reliability, passing 98% of propaganda and censorship tests. Available immediately via Devin Desktop and CLI (with rolling deployments on Web and Fusion), SWE-2 is designed to streamline software development cycles by delivering high-performance AI at an accessible price point without compromising trustworthiness or stability.






















