X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

By Dongjie Fu · Paper · cs.LG

While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data. To bridge this gap, we propose X$^3$-O

Cs.lg

View original

HomeResourceLoading…