# Jack Lindsey ## 概要 [[Anthropic]] の解釈可能性研究者。「Verbalizable Representations Form a Global Workspace in Language Models」(2026-07-06, Transformer Circuits Thread)の core contributor・correspondence author([email protected])。プロジェクトを統括した。(Source: [[@2026__TransformerCircuits__Verbalizable Representations Form a Global Workspace in Language Models]]) ## 貢献 - [[Wes Gurnee]] と共に、Jacobian Lens 手法および「言語化可能な表現とアクセス意識の接続」というアイデアを着想。 - J-space への directed modulation 実験、post-training が J-lens 読み出しに与える効果の初期実験を実施。 - [[Nicholas Sofroniew]] と共に、J-space をグローバルワークスペース理論の諸性質へ接続する実験を提案。 - flexible-vs-automatic タスク実験、J-space アブレーションが経験的報告(experiential reports)に与える効果の実験、脅迫シナリオ・Opus 4.6 のアラインメント監査事例・2 種のモデルオーガニズム実験(報酬ハッキング・報酬モデル迎合)を担当。 - counterfactual reflection training の手法を初期開発。 - [[Wes Gurnee]]・[[Nicholas Sofroniew]] と共に論文執筆を主導し、プロジェクトを監督した。 ## 出典 - [[@2026__TransformerCircuits__Verbalizable Representations Form a Global Workspace in Language Models]] — Author Contributions 節に基づく