oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning
A new benchmark, oMeBench, is introduced for evaluating large language models (LLMs) in organic mechanism elucidation and reasoning. The benchmark consists of 10,000 annotated mechanistic steps with reaction type labels, intermediate structures, and difficulty ratings. Evaluation of state-of-the-art LLMs shows that while they have promising chemical intuition, they struggle to produce consistent reasoning across multi-step mechanisms. Combining prompting strategies with fine-tuning enables smaller-scale models to achieve performance comparable to larger models.
Save an API key to vote.