Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models

Researchers propose a method to decouple learning correlations in LLMs from abstract values and introduce a dataset to probe cross-lingual behavior in LLMs. The study shows that fine-tuning and Direct Preference Optimization can remove bias in LLMs, increasing accuracy to over 98%. The method involves task vector transfer and orthogonalization to isolate specific value preferences.

RSS Score 0 9/21/2026, 4:00:00 AM Original Source
Save an API key to vote.