Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models
Researchers propose a method to decouple learning correlations in LLMs from abstract values and introduce a dataset to probe cross-lingual behavior in LLMs. The study shows that fine-tuning and Direct Preference Optimization can remove bias in LLMs, increasing accuracy to over 98%. The method involves task vector transfer and orthogonalization to isolate specific value preferences.
Save an API key to vote.