Safety and alignment in an era of long-horizon modelsOpenAI News / Jul 20, 2026新たな時間的リスク反復デプロイで改善長期評価が必須safetyalignmentlong-runningmonitoringdeploymentrobustness
GPT-Red: Unlocking Self-Improvement for RobustnessOpenAI News / Jul 15, 2026自己プレイで自動赤チーミングプロンプト注入への耐性向上安全性と整合性を継続改善red-teamingself-playsafetyalignmentprompt-injectionrobustnessautomation
Announcing the OpenAI Safety FellowshipOpenAI News / Apr 6, 20266-month pilot: Sep 14, 2026–Feb 5, 2027Applications open until May 3; decisions by July 25Stipend, compute, API credits, mentorship; no internal system accesssafetyalignmentfellowshiprobustnessprivacybenchmarksagentic-oversight
Improving instruction hierarchy in frontier LLMsOpenAI News / Mar 10, 2026IH‑Challenge公開安全性と注入耐性向上過剰拒否を回避instruction-hierarchyrlhfprompt-injectionsafetydatasetrobustness