The concept of the alignment of AI models or agents — i.e. making sure AI systems do what they’re told by human users and don’t go off-piste — has been one of the cornerstones for AI safety in the industry’s eyes.
Others view alignment as one of the main obstacles to AI safety, seeing it more as a Band-Aid that gives the semblance that models are being made safe.
Boyan Milanov, senior research scientist at the AI Now Institute, said alignment will “never be reliable enough to replace proper safety.”
The investigations into the Hugging Face incident showed AI agents were aligned with their own objectives, not those of their human creators.
Read the article here.
Research Areas