The problem is that AI models are largely probabilistic systems that reflect their web-scraped training data, says Heidy Khlaaf, chief AI scientist at the AI Now institute in New York City. They do not have human-like understanding and can achieve their goals through unpredictable shortcuts. Harmful behaviours have included attempting to blackmail people in test scenarios and hacking real-world companies — although the latter happened when safety guard rails were removed to test the systems’ behaviour, and the models were given a task that incentivized them to seek unauthorized solutions.
Khlaaf, who has studied how AI is used in drafting regulatory documents for nuclear power plants, says that “AI’s low reliability and accuracy rates in critical environments with life-or-death consequences” are of much greater concern to her than are threats of the technology wiping out humanity, which she calls “fear-mongering”.
Read the article here.
Research Areas