Source: Liberalism.org
by Rachel Lomasky
“The AI alignment problem asks whether the autonomous agents derived from Large Language Models and other AI technologies can reliably internalize the objectives for which they are optimized. The problem is magnified because human objectives are complex and hard to formalize, it is not possible to monitor the internals of the systems, and previous behavior is not necessarily a reliable predictor of how the AI will behave in novel situations. Worries that AIs might find goals that ‘score well’ but don’t match human intent have become a fashionable topic, driven in part by news stories about AI tools behaving in unexpected and disturbing ways, which can be difficult to monitor and reverse. … Human alignment, by comparison, has never been effectively enforced, although significant effort is spent on trying to determine our fellow humans’ motives.” (08/12/26)