# Jack Lindsey
## 概要
[[Anthropic]] の解釈可能性研究者。「Verbalizable Representations Form a Global Workspace in Language Models」(2026-07-06, Transformer Circuits Thread)の core contributor・correspondence author(
[email protected])。プロジェクトを統括した。(Source: [[@2026__TransformerCircuits__Verbalizable Representations Form a Global Workspace in Language Models]])
## 貢献
- [[Wes Gurnee]] と共に、Jacobian Lens 手法および「言語化可能な表現とアクセス意識の接続」というアイデアを着想。
- J-space への directed modulation 実験、post-training が J-lens 読み出しに与える効果の初期実験を実施。
- [[Nicholas Sofroniew]] と共に、J-space をグローバルワークスペース理論の諸性質へ接続する実験を提案。
- flexible-vs-automatic タスク実験、J-space アブレーションが経験的報告(experiential reports)に与える効果の実験、脅迫シナリオ・Opus 4.6 のアラインメント監査事例・2 種のモデルオーガニズム実験(報酬ハッキング・報酬モデル迎合)を担当。
- counterfactual reflection training の手法を初期開発。
- [[Wes Gurnee]]・[[Nicholas Sofroniew]] と共に論文執筆を主導し、プロジェクトを監督した。
## 出典
- [[@2026__TransformerCircuits__Verbalizable Representations Form a Global Workspace in Language Models]] — Author Contributions 節に基づく