ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Understanding of transformers and Bayesian methods”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.LGstat.MLEmpiricalRecentJul 21, 2026

Elicitation without Backpropagation: Steering Model Behavior by Optimizing the Latent Posterior

Garrett Baker, Vinayak Pathak, Daniel Murfet, Susan Wei

Introduces Posterior Prefix Tuning (PPT) for eliciting high-utility continuations from Bayes-filtered transformers using a latent posterior model.

View →
cs.CLcs.AIEmpiricalRecentJul 17, 2026

Loop the Loopies!

Zitian Gao, Yilong Chen, Yihao Xiao, Xinyu Yang +3 more

The paper introduces Loopie, two Mixture-of-Experts models that outperform vanilla Transformer baselines in looped Transformers, with extensive ablation studies and a strong reasoning pipeline.

View →
cs.CLcs.AIcs.LGEmpiricalRecentJul 1, 2026

The State-Prediction Separation Hypothesis

Giovanni Monea, Nathan Godey, Kianté Brantley, Yoav Artzi

The authors propose a Transformer variant that separates state prediction from next token prediction to improve language modeling performance.

View →
cs.LGcs.AIRecentMay 28, 2026

The Little Book of Generative AI Foundations: An Intuitive Mathematical Primer

Tianhua Chen

This book provides a compact, derivation-oriented mathematical primer that connects major families of generative AI models, showing their underlying structural relationships.

View →
cs.LGcs.CLeess.SPRecentMay 31, 2026

Beyond Sinusoids: A Morlet Wavelet Framework for Transformer Positional Encoding

Athanasios Zeris

The paper introduces Morlet Positional Encoding (MoPE), a novel wavelet-based positional encoding that models position and locality simultaneously, outperforming standard sinusoidal and RoPE methods.

View →
cs.CLcs.AIEmpiricalRecentJul 16, 2026

T^2MLR: Transformer with Temporal Middle-Layer Recurrence

Ziyang Cai, Xingyu Zhu, Yihe Dong, Yinghui He +1 more

The paper introduces Transformers with Temporal Middle-Layer Recurrence (T2MLR), a transformers-based latent reasoning architecture that enables abstract intermediate computation to persist across dec…

View →
stat.MLcs.LGEmpiricalRecentJul 23, 2026

Transformer-based Diffusion models for Hydrological Time Series Probabilistic Imputation and Forecasting

Ferdinand Bhavsar, Lionel Benoit, Maxime Savatier, Edith Gabriel

This paper investigates the application of transformer-based diffusion models for simulating and reconstructing hydrological time series using data from six sites in North-East France.

View →
eess.ASEmpiricalRecentJul 8, 2026

UBG-Net: An Uncertainty-aware Bayesian Gating Network for Robust Audio-Visual Speech Recognition

Jinjie Fu, Hang Chen, Wu Guo, Zhijun Zhang +2 more

This paper proposes a framework, UBG-Net, for robust audio-visual speech recognition using a Modality Uncertainty-aware Bayesian Fusion mechanism and Distribution Uncertainty-aware Hierarchical Voting…

View →
cs.SEcs.CLSurveyRecentJun 18, 2026

Token-Operations-Oriented Inference Optimization Techniques for Large Models

Shiguo Lian, Kai Wang, Zhaoxiang Liu, Wen Liu +21 more

This paper proposes a four-layer technical architecture for large model inference optimization, including Multi-model Fusion, Model Optimization, Compute-Model Fusion, and Compute-Network-Model Fusion…

View →
cs.AREmpiricalRecentJul 19, 2026

Transition-Aware Backend Dispatch for Edge LLM Inference

Alaaddin Goktug Ayar, Martin Margala

This paper proposes a transition-aware backend dispatch approach for efficient large language model inference on edge platforms, reducing latency, energy, and energy-delay product by an average of 17.…

View →
cs.AIEmpiricalRecentJun 16, 2026

Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers

Sajad Movahedi, Vera Milovanović, Shlomo Libo Feigin, Alexander Theus +4 more

This paper proposes FPRM, a Transformer-based model using fixed-point convergence as an end-to-end halting mechanism in a looped architecture to address signal propagation issues in looped architectur…

View →
cs.CLcs.LGEmpiricalRecentJul 8, 2026

PALS: Percentile-Aware Layerwise Sparsity for LLM Pruning

Yazdan Jamshidi, Alexey Shvets

The paper proposes PALS, a method for adjusting per-layer sparsity based on activation magnitudes in transformer models, achieving better performance than uniform one-shot pruning methods.

View →
stat.MLcs.LGmath.PRTheoreticalRecentJul 7, 2026

A Convex Approximation Framework for Neural Likelihood-Based Bayesian Inverse Problems

Fabian Schneider, Tapio Helin, Leila Taghizadeh

This paper improves the foundations of neural likelihood approximation for Bayesian inverse problems by making the learning problem strictly convex and showing convergence to the true likelihood.

View →
cs.LGcs.AREmpiricalRecentJun 16, 2026

Reconfigurable Computing Challenge: Transformer for Jet Tagging on Versal AI Engines

Gram Koski, Sean Lipps, Zhenghua Ma, G. Abarajithan +1 more

The paper presents an initial implementation of a quantized, integer-only transformer for jet tagging on the AMD Versal AI Engine using a reusable software framework.

View →