نسخة أولية وصول مفتوح
HARPO: Hallucination-Aware Reinforcement Learning for Faithful and Creative Language Generation
Large Language Models (LLMs) are prone to generating hallucinated content, which compromises their reliability in knowledge-intensive tasks. To address this challenge without sacrificing creativity, we propose HARPO, a reinforcement learning framework designed to jointly optimize faithfulness and creativity. HARPO inco …