Abstention vs. Hallucination: Benchmarking LLM Source Attribution for Scientific Citations
A new benchmark, REASONS, is introduced to evaluate the reliability of scientific citation attribution in large language models (LLMs). The benchmark assesses LLMs' ability to correctly attribute citations and abstain from providing incorrect information when uncertain. The evaluation framework includes various settings, such as zero-context and retrieval-augmented prompting, to test LLMs' performance. The results show that certain configurations can significantly improve citation attribution accuracy, but also highlight the challenges of balancing reliability and responsiveness.
Save an API key to vote.