Research1 min read
HarvestBench measures LLM agent decisions on animal harm in farm simulation
HarvestBench evaluates how language models decide to avoid harming animals in a simulated farm environment, with decisions priced and measured across multiple models and scenarios.
From arXiv cs.AI



