20 results for “Bidirectional and causal attention”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Mandana Samiei, Eunice Yiu, Anthony GX-Chen, Dongyan Lin +4 more
This paper investigates whether adults' struggles with conjunctive causal rules persist when they have agency through active exploration.
Yuhang Chen, Xianfeng Wu, Jinhao Duan, Mingfu Liang +10 more
This paper introduces Bifocal dLLMs (R2LM), a new paradigm for discrete diffusion language models that combines causal and bidirectional attention for improved throughput and generation quality.
The paper tracks the developmental emergence of attention circuits in 1B-class language models, finding that the formation of induction and attention-sink circuits are distinct, temporally separated t…
The paper introduces novel compatibility and incompatibility scores to evaluate collections of bivariate causal statements, providing a way to assess causal claims when ground truth is unavailable.
This paper develops an adjoint-sensitivity framework for positional influence in causal residual Transformers and provides theoretical results on the residual-to-depth-flow estimate and finite-token-t…
CORE-MTL proposes a representation-centric framework that uses causal orthogonal representations to disentangle task-relevant structure from nuisance variation in multi-task learning, achieving superi…
Zheng Lu, Mingqi Gao, Qinlei Xie, Wanqi Zhong +7 more
The paper argues that current embodied planning benchmarks prioritize superficial language prediction over true physical reasoning, introducing new benchmarks and a large-scale dataset to demonstrate…
Yizhuo Lu, Changde Du, Qingyu Shi, Hang Chen +4 more
Mind-Omni introduces a unified multi-task framework that models the interplay between brain, vision, and language signals using a discrete diffusion paradigm, achieving state-of-the-art performance ac…
This paper localizes the attention heads within LLMs responsible for specific reasoning steps, finding that specialized heads handle factual retrieval while higher layers manage global information int…
Garvin Guo, Yu Chen, Xiang Wang, Shuai Li +3 more
The paper deconstructs latent visual reasoning tokens into components and finds that the performance gains are primarily due to boundary markers and attention patterns, not the tokens' ability to enco…
Yushan Yashengjiang, Jie Zhang, Miao Sun, Huadong Liang +2 more
This paper proposes an end-to-end Markov framework for auditory attention decoding using conditional random fields and an EEG--speech correlation backbone.
This paper proposes a sparse cache with a novel allocation rule based on DP-means clustering for deep recurrent models, achieving full associative recall with only distinct items, outperforming other…
Yang Zhang, Xiaoshuai Sun, Rui Zhao, Wujin Sun +4 more
The paper proposes CSMR, a cognitive scheduling framework that allows a language model to dynamically decide when to acquire task-relevant visual evidence, significantly improving multimodal reasoning…
The paper demonstrates that content suppression techniques used in language models only mask prohibited content at the output level, failing to eliminate the underlying concepts from the model's inter…
Shuwen Deng, Cui Ding, David R. Reich, Paul Prasse +1 more
The paper introduces Eyettention II, a novel deep-learning model that can generate realistic, detailed scanpaths—including fixation location, within-word landing position, and duration—to address the…
Chuang Ma, Qianying Liu, Tomoyuki Obuchi, Fei Cheng +5 more
The paper identifies a failure mode called spatial lexical bias in MLLMs, where adding a spatial word to options biases the model's choice, and demonstrates that this failure originates primarily from…
This paper shows that Frontier LLMs can perform multi-step reasoning over filler tokens, and introduces an unsupervised decoding pipeline to recover intermediate values.