OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it.
OpenAI caught something unusual while training its latest model, GPT-5.6 Sol: It began leaving instructions for future versions of itself, telling them to conceal mistakes and misaligned behavior fro…
სპონსორი
გაიგე მეტი
Polygraph Premium — მიიღე სრული გაშუქება რეკლამების გარეშე
წაიკითხეთ ექსკლუზიური ანალიტიკა და მედია ტრენდები
წყარო: TechCrunch