Architectures
Neural network architectures, attention mechanisms, and model design
20 papers indexed
Forget Attention: Importance-Aware Attention Is All You Need
The paper proposes SISA (SSM-Informed Softmax Attention), a novel hybrid attention mechanism that integrates state-space model (SSM) importance signals directly into the attention score, achieving sta…
Zamba2-VL Technical Report
Zamba2-VL is a new suite of vision-language models built on the Zamba2 hybrid architecture, achieving state-of-the-art performance and significantly improved inference efficiency compared to leading T…
Bifocal Diffusion Language Models: Asymmetric Bidirectional Context for Parallel Generation
Yuhang Chen, Xianfeng Wu, Jinhao Duan, Mingfu Liang +10 more
This paper introduces Bifocal dLLMs (R2LM), a new paradigm for discrete diffusion language models that combines causal and bidirectional attention for improved throughput and generation quality.
AEGIS: Adversarial Entropy-Guided Immune System -- Thermodynamic State Space Models for Zero-Day Network Evasion Detection
AEGIS introduces a novel physics-based system that analyzes encrypted network traffic flow dynamics, achieving state-of-the-art zero-day evasion detection with high accuracy and low latency.
Task Structure Reverses Layerwise State Encoding in Sequence Models
The paper demonstrates that the location and nature of state encoding in sequence models are not fixed architectural traits but are highly dependent on the specific task, showing that the encoding pro…
CaMBRAIN: Real-time, Continuous EEG Inference with Causal State Space Models
CaMBRAIN introduces a novel Mamba-based State Space Model (SSM) for real-time, continuous EEG inference, achieving state-of-the-art results with significantly higher throughput than existing methods.
IR275K: A Benchmark for Infrared Multi-Frame Super-Resolution Toward Efficient Remote Sensing
Jie Deng, Heyang Wang, Changxin Wang, Junkai Shen +5 more
This paper introduces IR275K, a curated benchmark for multi-frame super-resolution in infrared remote sensing, and evaluates CGMamba, a lightweight state-space model, achieving state-of-the-art perfor…
MARS: Multi-rate Aggregation of Recency Signals for Sequential Recommendation across Sparse and Dense Regimes
MARS proposes an encoder-agnostic aggregation operator that explicitly models multi-scale temporal structure in sequential recommendation, achieving state-of-the-art performance across both sparse and…
EVOTS: Evolutionary Transformer Search for Time Series Forecasting
This paper introduces EVOTS, an evolutionary neural architecture search framework for discovering task-adaptive Transformer-like models for multivariate time-series forecasting.
Learning the Signature of Memorization in Autoregressive Language Models
The paper introduces a novel, transferable learned attack (LT-MIA) that detects a universal 'signature of memorization' in language models, achieving high accuracy across diverse model architectures (…
Safety, Security, and Cognitive Risks in State-Space Models: A Systematic Threat Analysis with Spectral, Stateful, and Capacity Attacks
This paper provides the first systematic threat analysis of State-Space Models (SSMs) in safety-critical applications, introducing novel attack classes and formal metrics to quantify their security an…
DAPGNet: Dynamic Adaptive Physics-Guided Graph Diffusion Network for Hyperspectral Image Classification
This paper proposes DAPGNet, a dynamic adaptive physics-guided graph diffusion network for hyperspectral image classification, which achieves state-of-the-art performance on four datasets.
Learning the Arabic Dialect Continuum as a Continuous Space: A Regression Approach to Speaker Origin Prediction
Mohamed Aziz Khadraoui, Adel Ammar, Bilel Benjdira, Zahid Khan +2 more
A regression-based approach is presented for Arabic dialect geolocation using a hierarchical neural architecture and spherical geodesic loss, achieving a median localization error of 481.2 km.
SE-Enhanced ViT and BiLSTM-Based Intrusion Detection for Secure IIoT and IoMT Environments
The paper proposes an SE ViT-BiLSTM hybrid model for enhanced intrusion detection in IIoT and IoMT environments, achieving superior performance on real-world datasets, especially after data balancing.
ConsisFormer: Compute-Efficient Transformer for Wireless Foundation Models Based on Channel Consistency
Yuwei Wang, Li Sun, Tingting Yang, Liwen Jing +3 more
This paper proposes ConsisFormer, a compute-efficient Transformer design for wireless foundation models (WFMs) using short-term channel consistency and adaptive token aggregation.
SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer
Yuyang Zhao, Yicheng Pan, Qiyuan He, Jincheng Yu +5 more
SANA-Streaming introduces a novel, efficient framework that enables real-time, high-resolution streaming video-to-video editing by combining a hybrid diffusion transformer with specialized training an…
EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction
EigeNet introduces a geometry-informed multi-modal Transformer framework to achieve state-of-the-art few-shot novel view Room Impulse Response (RIR) prediction by effectively integrating spatial geome…
The Giant Hippocampus: From Structural Monoculture to a System of Systems
This paper argues for the importance of modularity and heterogeneity in AI architectures, contrasting the Transformer model with the structure of the cortex.
Detection vs. Execution: Single-Bucket Probes Miss Half the Mamba-2 State Sink
The paper demonstrates that in Mamba-2, single-bucket probes can detect a large functional signature (detection layer) that is not fully responsible for the actual computation (execution layer), chall…
ITP-STDP: An Intrinsic-Timing Power-of-Two Learning Engine for On-Chip SNN Training
Haihang Xia, Xinyu Zhao, Xuecheng Wang, John Goodenough +4 more
This paper proposes and validates a novel hardware architecture, ITP-STDP, to significantly reduce the energy consumption and hardware overhead associated with training Spiking Neural Networks (SNNs).