Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:
ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Home/Authors/Yimin

Yimin

46 indexed papers

Recent (6 mo)
46
With code
0
Influential cites
0
Benchmarked
0

Publications per year

46
26

Top categories

AI×32Crypto×16ML×15Vision×9NLP×8Software Eng.×4Distributed×3Multimedia×3

Frequent co-authors

Yiming Zhang7×
Yiming Li5×
Yiming Liu4×
Yiming Wang3×
Wei Zhou2×
Koji Tsuda2×

Research Timeline

2026
Implicit Drifting Policy: One-Step Action Generation via Conditional Expert Geometry

The Implicit Drifting Policy (IDP) is a novel one-step action generation framework that implicitly enforces trajectory correction constraints by analyzing local expert action geometry, overcoming the difficulties of explicitly estimating a training-time drifting field.

Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning

The paper introduces Med-HEAL, a comprehensive framework and dataset for systematically identifying and mitigating hallucinations in medical LLMs, demonstrating that a self-critique pipeline significantly improves model accuracy.

SimSD: Simple Speculative Decoding in Diffusion Language Models

The paper proposes SimSD, a plug-and-play speculative decoding algorithm that adapts diffusion language models (dLLMs) to achieve fast, token-level acceleration by restoring causal masking capabilities.

Order within Chaos: Capturing Intrinsic Energy Anomalies for AI-Manipulated Image Forgery Localization

The paper proposes FLAME, a novel framework that detects AI-generated image forgeries by identifying intrinsic energy anomalies caused by the diffusion process, achieving state-of-the-art localization.

CAPF: Guiding Search-Agent Rollouts with Credit-Attenuated Privileged Feedback

The paper proposes Credit-Attenuated Privileged Feedback (CAPF), a training-time mechanism that uses verifier-side information to guide LLM search agents, significantly improving their performance on complex QA tasks.

Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners

The paper argues that current embodied planning benchmarks prioritize superficial language prediction over true physical reasoning, introducing new benchmarks and a large-scale dataset to demonstrate that physically grounded causal reasoning is necessary for reliable autonomous agents.

Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models

This paper introduces Imaginative Perception Tokens (IPT) to improve spatial reasoning in vision language models.

ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues

The paper introduces ReproRepo, a scalable framework for evaluating the reproducibility of machine learning research using LLM agents and human-raised GitHub issues.

Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation

DigenRL is a disaggregated RL framework for diffusion-based generative LLMs that achieves 1.56-2.10x throughput improvements over state-of-the-art diffusion RL systems.

DiStash: A Disaggregated Multi-Stash Transactional Key-Value Store

This paper introduces DiStash, a disaggregated transactional key-value store that enables an application to use a single transaction to manage key-value pairs across different pools of stashes, preventing race conditions and data loss.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

This paper proposes Joint Speech-Text Interleaved Pretraining (JSTIP) for speech recognition, which constructs interleaved speech-text sequences and achieves consistent entity accuracy improvement.

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

The paper introduces MedPMC, a framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models, resulting in improved performance on various benchmarks.

DiPhon: Diffusion on Graphons for Scalable Graph Generation

This paper introduces DiPhon, a diffusion framework for size-scalable graph generation, using a continuous diffusion process on the graphon space and a discretized graph-level process.

Native Video-Action Pretraining for Generalizable Robot Control

This paper introduces LingBot-VA 2.0, a video-action foundation model designed for embodiment, with semantic visual-action tokenization, causal pretraining, sparse MoE backbone, and enhanced asynchronous inference.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

The paper introduces PS4, a framework for training target speaker extraction models using a large-scale corpus and proxy-supervised joint training strategy.

Scalable Visual Pretraining for Language Intelligence

This paper presents the benefits of visual pretraining for foundation model intelligence, outperforming text-only pretraining on multiple backbones and benchmarks.

SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation

The paper proposes SAGA, a framework for schema-aware grounding in agentic text-to-SPARQL generation, which maintains a persistent type state, filters incompatible property candidates, and handles missing schema information permissively.

Scaling Unmodified Multithreaded Applications with Elastic CXL-based Distributed Shared Memory

xDSM is a full-space, elastic DSM system built over CXL that transparently scales unmodified multithreaded applications by employing an OS-runtime co-design, dynamic data placement policy, and spatial locality-aware elasticity.

Climate-resilient electric vehicle charging infrastructure for sustainable cities: An interpretable causal-ensemble framework for preventive maintenance and low-carbon mobility

This paper develops FGDSE, a feature-governed dynamic stacking ensemble for climate-resilient charging-asset management in electric vehicles, which predicts daily fault risk over a multi-week horizon and identifies extreme heat as a causal factor for heat-sensitive posts.

Specula: Scaling formal specifications for autonomous model checking of system code

Specula is an autonomous system that generates high-quality formal specifications for large, complex code using LLMs, improving understanding and finding bugs.

Highlighted terms show continued research focus across papers

Papers

cs.SEcs.AIcs.DCNEWEmpiricalJul 28, 2026

Specula: Scaling formal specifications for autonomous model checking of system code

Qian Cheng, Saad Mohammad Rafid Pial, Ruize Tang, Yiming Su +5 more

Specula is an autonomous system that generates high-quality formal specifications for large, complex code using LLMs, improving understanding and finding bugs.

View →
math.OCcs.LG
Empirical
Recent
Jul 23, 2026

Climate-resilient electric vehicle charging infrastructure for sustainable cities: An interpretable causal-ensemble framework for preventive maintenance and low-carbon mobility

Cande Lian, Wentao Zeng, Jiabin Wu, Yiming Bie +1 more

This paper develops FGDSE, a feature-governed dynamic stacking ensemble for climate-resilient charging-asset management in electric vehicles, which predicts daily fault risk over a multi-week horizon…

View →
cs.OSEmpiricalRecentJul 17, 2026

Scaling Unmodified Multithreaded Applications with Elastic CXL-based Distributed Shared Memory

Guowei Liu, Kang Chen, Laiping Zhao, Yiming Li +6 more

xDSM is a full-space, elastic DSM system built over CXL that transparently scales unmodified multithreaded applications by employing an OS-runtime co-design, dynamic data placement policy, and spatial…

View →
cs.AIcs.IRcs.LGEmpiricalRecentJul 16, 2026

SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation

Yiming Zhang, Koji Tsuda

The paper proposes SAGA, a framework for schema-aware grounding in agentic text-to-SPARQL generation, which maintains a persistent type state, filters incompatible property candidates, and handles mis…

View →
cs.CVcs.AIcs.MMEmpiricalRecentJul 10, 2026

Scalable Visual Pretraining for Language Intelligence

Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao +12 more

This paper presents the benefits of visual pretraining for foundation model intelligence, outperforming text-only pretraining on multiple backbones and benchmarks.

View →
cs.ROcs.CVEmpiricalRecentJul 9, 2026

Native Video-Action Pretraining for Generalizable Robot Control

Qihang Zhang, Lin Li, Luyao Zhang, Shuai Yang +25 more

This paper introduces LingBot-VA 2.0, a video-action foundation model designed for embodiment, with semantic visual-action tokenization, causal pretraining, sparse MoE backbone, and enhanced asynchron…

View →
cs.SDcs.AIEmpiricalRecentJul 9, 2026

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

Wanyi Ning, Wei Zhou, Yingpeng Li, Yinshang Guo +2 more

The paper introduces PS4, a framework for training target speaker extraction models using a large-scale corpus and proxy-supervised joint training strategy.

View →
cs.CVcs.LGEmpiricalRecentJul 8, 2026

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

Hyunjae Kim, Dain Kim, Pan Xiao, Serina S. Applebaum +24 more

The paper introduces MedPMC, a framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models, resulting in improved performance on various…

View →
stat.MLcs.AIcs.LGTheoreticalRecentJul 8, 2026

DiPhon: Diffusion on Graphons for Scalable Graph Generation

Sergio Rozada, Yiming Qin, Manuel Madeira, Pascal Frossard +1 more

This paper introduces DiPhon, a diffusion framework for size-scalable graph generation, using a continuous diffusion process on the graphon space and a discretized graph-level process.

View →
cs.CLeess.ASEmpiricalRecentJul 2, 2026

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

Ruchao Fan, Yiming Wang, Rui Zhao, Liliang Ren +9 more

This paper proposes Joint Speech-Text Interleaved Pretraining (JSTIP) for speech recognition, which constructs interleaved speech-text sequences and achieves consistent entity accuracy improvement.

View →
cs.DBcs.DCcs.PFEmpiricalRecentJun 26, 2026

DiStash: A Disaggregated Multi-Stash Transactional Key-Value Store

Yiming Gao, Hieu Nguyen, Jun Li, Shahram Ghandeharizadeh

This paper introduces DiStash, a disaggregated transactional key-value store that enables an application to use a single transaction to manage key-value pairs across different pools of stashes, preven…

View →
cs.AIcs.DCcs.NIEmpiricalRecentJun 23, 2026

Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation

Sijie Wang, Zhengyu Qing, Zhiqiang Tan, Yiming Yin +5 more

DigenRL is a disaggregated RL framework for diffusion-based generative LLMs that achieves 1.56-2.10x throughput improvements over state-of-the-art diffusion RL systems.

View →
cs.CLcs.AIcs.LGEmpiricalRecentJun 16, 2026

ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues

Shanda Li, Qiuhong Anna Wei, Jingwu Tang, Valerie Chen +4 more

The paper introduces ReproRepo, a scalable framework for evaluating the reproducibility of machine learning research using LLM agents and human-raised GitHub issues.

View →
cs.AIRecentJun 2, 2026

Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models

Mahtab Bigverdi, Lindsey Li, Weikai Huang, Yiming Liu +7 more

This paper introduces Imaginative Perception Tokens (IPT) to improve spatial reasoning in vision language models.

View →
cs.CLcs.AIRecentJun 1, 2026

SimSD: Simple Speculative Decoding in Diffusion Language Models

Junxia Cui, Haotian Ye, Runchu Tian, Hongcan Guo +8 more

The paper proposes SimSD, a plug-and-play speculative decoding algorithm that adapts diffusion language models (dLLMs) to achieve fast, token-level acceleration by restoring causal masking capabilitie…

View →
cs.CVcs.AIRecentJun 1, 2026

Order within Chaos: Capturing Intrinsic Energy Anomalies for AI-Manipulated Image Forgery Localization

Yiming Wang, Baiqi Wu, Qingming Li, Jiahao Chen +2 more

The paper proposes FLAME, a novel framework that detects AI-generated image forgeries by identifying intrinsic energy anomalies caused by the diffusion process, achieving state-of-the-art localization…

View →
cs.AIRecentJun 1, 2026

CAPF: Guiding Search-Agent Rollouts with Credit-Attenuated Privileged Feedback

Bin Chen, Xinye Liao, Yiming Liu, Xin Liao +1 more

The paper proposes Credit-Attenuated Privileged Feedback (CAPF), a training-time mechanism that uses verifier-side information to guide LLM search agents, significantly improving their performance on…

View →
cs.AIRecentJun 1, 2026

Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners

Zheng Lu, Mingqi Gao, Qinlei Xie, Wanqi Zhong +7 more

The paper argues that current embodied planning benchmarks prioritize superficial language prediction over true physical reasoning, introducing new benchmarks and a large-scale dataset to demonstrate…

View →
cs.ROcs.AIRecentMay 31, 2026

Implicit Drifting Policy: One-Step Action Generation via Conditional Expert Geometry

Zemin Yang, Yaoyu He, Yiming Zhong, Yuhao Zhang +4 more

The Implicit Drifting Policy (IDP) is a novel one-step action generation framework that implicitly enforces trajectory correction constraints by analyzing local expert action geometry, overcoming the…

View →
cs.CLRecentMay 31, 2026

Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning

Yiming Liao, Zeno Franco, Jose Eduardo Lizarraga Mazaba, Keke Chen

The paper introduces Med-HEAL, a comprehensive framework and dataset for systematically identifying and mitigating hallucinations in medical LLMs, demonstrating that a self-critique pipeline significa…

View →