ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “DiT model”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CVEmpiricalRecentJul 8, 2026

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang +23 more

This paper introduces LingBot-Video, a video pretraining paradigm for embodied intelligence using a DiT-based approach, Mixture-of-Experts framework, and extensive robot-oriented data.

View →
cs.LGcs.CVEmpiricalRecentJul 3, 2026

When Geometry Aligns: Dihedral Hidden-State Transformations in UNet, ViT, and DiT Architectures

Mojtaba Faramarzi, Alex Lamb, Irina Rish

This paper explores the effects of geometric perturbations on the stability of diffusion architectures, specifically Diffusion Transformers and Convolutional UNets, using a unified framework and contr…

View →
cs.CVcs.AIcs.LGRecentMay 30, 2026

DASH: Dual-Branch Score Distillation for Guidance-Calibrated Compact Diffusion Models

Abdullah Al Shafi, Kazi Saeed Alam, Sk Imran Hossain, Engelbert Mephu Nguifo

DASH introduces a dual-branch distillation framework to effectively compress class-conditional diffusion models by independently supervising both score branches, significantly preserving guidance fide…

View →
cs.LGcs.AIcs.SDRecentMay 30, 2026

Logit Distillation on Manifolds: Mapping by Learning

Yiru Yang, Junling Wang, Nishant Kumar Singh, Luohong Wu +1 more

The paper proposes a novel layer and point-wise projection mapping combined with LoRA injection to efficiently distill knowledge from a large teacher model to a small student model, significantly impr…

View →
cs.CRRecentApr 10, 2026

Hagenberg Risk Management Process (Part 3): Operationalization, Probabilities, and Causal Analysis

Eckehard Hermann, Harald Lampesberger

The paper introduces a comprehensive framework, Realtime Risk Studio, that operationalizes qualitative risk models (Bowtie diagrams) into formal, probabilistic, and intervention-ready runtime models u…

View →
cs.AIcs.ARcs.CEEmpiricalRecentJul 7, 2026

Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation

Niels Potters, Theo Hofman

A black-box evaluation framework is presented to assess Large Language Models' ability to generate Design Structure Matrices from technical documentation, using structural, classification, and stabili…

View →
cs.CRcs.AIcs.CLRecentApr 7, 2026

Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts

Fatih Uenal

This paper introduces Swiss-Bench 003, an expanded evaluation framework assessing LLM reliability and adversarial security across eight dimensions using 808 Swiss-specific items, revealing that self-g…

View →
cs.DBcs.DCcs.PFEmpiricalRecentJun 26, 2026

DiStash: A Disaggregated Multi-Stash Transactional Key-Value Store

Yiming Gao, Hieu Nguyen, Jun Li, Shahram Ghandeharizadeh

This paper introduces DiStash, a disaggregated transactional key-value store that enables an application to use a single transaction to manage key-value pairs across different pools of stashes, preven…

View →
cs.CEmath.NARecentMay 29, 2026

A non-intrusive approach to index-aware learning

Peter Förster, Idoia Cortes Garcia, Wil Schilders, Sebastian Schöps

The paper introduces a non-intrusive variant of index-aware learning for solving differential-algebraic equations (DAEs), ensuring that learned solutions maintain physical consistency like Kirchhoff's…

View →
cs.LGmath.DGmath.OCEmpiricalRecentJun 28, 2026

Dead-Direction Conditioners: Gauge-Equivariant Preconditioning for Deep Networks

Tejas Pradeep Shirodkar

This paper introduces DDC, a Dead-Direction Conditioner that keeps a deep network's optimization on the symmetry quotient by conditioning the optimizer's state in the orbit decomposition of a $G$-inva…

View →
cs.LGcs.CRRecentApr 16, 2026

FedIDM: Achieving Fast and Stable Convergence in Byzantine Federated Learning through Iterative Distribution Matching

He Yang, Dongyi Lv, Wei Xi, Song Ma +2 more

FedIDM introduces a novel federated learning framework that uses iterative distribution matching to achieve fast and stable convergence and maintain high model utility even when facing a large proport…

View →
stat.MLcs.LGmath.OCTheoreticalRecentJul 20, 2026

An Adjoint-Sensitivity Framework for Lost-in-the-Middle Phenomena in Causal Residual Transformers

Cheng Huan, Hongwei Yuan

This paper develops an adjoint-sensitivity framework for positional influence in causal residual Transformers and provides theoretical results on the residual-to-depth-flow estimate and finite-token-t…

View →
cs.LGcs.CLEmpiricalRecentJul 16, 2026

On-Policy Delta Distillation

Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han

This paper introduces a new distillation reward, the delta signal, for on-policy distillation in reinforcement learning, which captures changes induced by reasoning tuning and improves performance.

View →
cs.LGcs.AIRecentMay 29, 2026

De-attribute to Forget for LLM Unlearning

Xinyang Lu, Jiabao Pan, Rachael Hwee Ling Sim, See-Kiong Ng +2 more

The paper proposes DareU, a novel LLM unlearning framework that optimizes unlearning by zeroing out data attribution scores instead of maximizing prediction loss, achieving effective unlearning while…

View →
cs.LGcs.AIRecentMay 28, 2026

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models

Xiaohang Tang, Keyue Jiang, Che Liu, Qifang Zhao +3 more

The paper proposes Guided Denoiser Self-Distillation (GDSD), a novel method that bypasses the use of likelihood surrogates (like ELBO) in RL for diffusion language models, achieving state-of-the-art p…

View →
cs.LGcs.AIRecentMay 28, 2026

LLMs Without Deep Neural Networks: New Architecture, Benefits and Case Study

Vincent Granville

The paper introduces a novel, non-deep neural network architecture that achieves the performance of LLMs by finding the global optimum of the loss function in a single, closed-form iteration, eliminat…

View →
cs.NEEmpiricalRecentJun 26, 2026

DE-2LS: Differential Evolution with Lightweight Late Local Search for Constrained Numerical Optimization

Dikshit Chauhan, Anupam Trivedi

This paper proposes DE-2LS, a late-stage, locally search-enhanced variant of differential evolution that improves the exploitation capability of the RDEx-based search framework while preserving its sp…

View →
cs.AIRecentMay 31, 2026

DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts

Jiarui Feng, Hanqing Zeng, Karish Grover, Ruizhong Qiu +10 more

The paper proposes DAG-MoE, a novel sparse Mixture-of-Experts framework that replaces standard weighted-sum aggregation with structural aggregation to enhance model performance and enable multi-step r…

View →