跳到正文
原文
arXiv cs.SD RSS·· 1 天前

音频 token 注意力在语言模型运行前即可预测

Audio Token Attention Is Predictable Before the Language Model Runs

摘要

研究发现,大型音频语言模型(LALM)中音频 token 将获得的注意力,可在语言模型运行前由其编码器输出线性预测,在 13 个 LALM 中的 11 个上全层注意力排序相关系数 ρ≥0.69。

应来源方要求,这里只提供摘要与原文入口。完整内容请阅读原文。

来源:arXiv cs.SD RSS · arxiv.org