20 results for “Understanding of transformers and Bayesian methods”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Introduces Posterior Prefix Tuning (PPT) for eliciting high-utility continuations from Bayes-filtered transformers using a latent posterior model.
Zitian Gao, Yilong Chen, Yihao Xiao, Xinyu Yang +3 more
The paper introduces Loopie, two Mixture-of-Experts models that outperform vanilla Transformer baselines in looped Transformers, with extensive ablation studies and a strong reasoning pipeline.
The authors propose a Transformer variant that separates state prediction from next token prediction to improve language modeling performance.
This book provides a compact, derivation-oriented mathematical primer that connects major families of generative AI models, showing their underlying structural relationships.
The paper introduces Morlet Positional Encoding (MoPE), a novel wavelet-based positional encoding that models position and locality simultaneously, outperforming standard sinusoidal and RoPE methods.
Ziyang Cai, Xingyu Zhu, Yihe Dong, Yinghui He +1 more
The paper introduces Transformers with Temporal Middle-Layer Recurrence (T2MLR), a transformers-based latent reasoning architecture that enables abstract intermediate computation to persist across dec…
This paper investigates the application of transformer-based diffusion models for simulating and reconstructing hydrological time series using data from six sites in North-East France.
Jinjie Fu, Hang Chen, Wu Guo, Zhijun Zhang +2 more
This paper proposes a framework, UBG-Net, for robust audio-visual speech recognition using a Modality Uncertainty-aware Bayesian Fusion mechanism and Distribution Uncertainty-aware Hierarchical Voting…
Shiguo Lian, Kai Wang, Zhaoxiang Liu, Wen Liu +21 more
This paper proposes a four-layer technical architecture for large model inference optimization, including Multi-model Fusion, Model Optimization, Compute-Model Fusion, and Compute-Network-Model Fusion…
This paper proposes a transition-aware backend dispatch approach for efficient large language model inference on edge platforms, reducing latency, energy, and energy-delay product by an average of 17.…
This paper proposes FPRM, a Transformer-based model using fixed-point convergence as an end-to-end halting mechanism in a looped architecture to address signal propagation issues in looped architectur…
The paper proposes PALS, a method for adjusting per-layer sparsity based on activation magnitudes in transformer models, achieving better performance than uniform one-shot pruning methods.
This paper improves the foundations of neural likelihood approximation for Bayesian inverse problems by making the learning problem strictly convex and showing convergence to the true likelihood.
Gram Koski, Sean Lipps, Zhenghua Ma, G. Abarajithan +1 more
The paper presents an initial implementation of a quantized, integer-only transformer for jet tagging on the AMD Versal AI Engine using a reusable software framework.