Some question whether a hybrid human-AI judge is the right solution. Sarah Myers West, co-executive director of the AI Now Institute, a policy research center, said the industry has relied too heavily on general benchmarks. She wants safety researchers to develop evaluations tailored to all the different ways real people use the technology in their daily lives. She noted that in medicine, for instance, a drug isn’t judged as safe across the board — it’s tested as a treatment for a specific illness, dose and type of patient.

“What is it that you’re benchmarking if you’re not talking about specific use cases?” she said.

Read the full article here.

Research Areas