Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:
ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Home/Authors/Kai Li

Kai Li

22 indexed papers

Recent (6 mo)
22
With code
0
Influential cites
0
Benchmarked
0

Publications per year

22
26

Top categories

AI×8Crypto×8Robotics×5Vision×4ML×4NLP×4Audio and Speech Processing×2Software Eng.×2

Frequent co-authors

Wenkai Li5×
Xiaoqi Li4×
Zongwei Li4×
Sikai Li3×
Zhenyu Wei3×
Yunchao Yao3×

Research Timeline

2026
Say Something Else: Rethinking Contextual Privacy as Information Sufficiency

The paper proposes Information Sufficiency (IS) as a comprehensive framework for privacy-preserving LLM communication, demonstrating that free-text pseudonymization outperforms existing suppression and generalization methods, especially in multi-turn conversations.

PSR2: A Phase-based Semantic Reasoning Framework for Atomicity Violation Detection via Contract Refinement

The paper introduces PSR extsuperscript{2}, a novel static analysis framework that significantly improves the detection of atomicity violations in smart contracts by combining structural path searching with deep semantic reasoning.

NFTDELTA: Detecting Permission Control Vulnerabilities in NFT Contracts through Multi-View Learning

NFTDELTA is a novel framework that uses multi-view learning on static code analysis to detect permission control vulnerabilities in NFT contracts with high accuracy.

SafeDream: Safety World Model for Proactive Early Jailbreak Detection

SAFEDREAM introduces a lightweight, external world-model framework that proactively detects multi-turn jailbreak attacks by modeling cumulative safety erosion and predicting early failure points.

Heimdallr: Characterizing and Detecting LLM-Induced Security Risks in GitHub CI Workflows

This paper introduces Heimdallr, a novel framework that characterizes and detects LLM-induced security risks by analyzing the full execution chain of LLM integrations within GitHub CI workflows.

Graph Representation Learning Augmented Model Manipulation on Federated Fine-Tuning of LLMs

The paper proposes an Augmented Model maniPulation (AugMP) strategy, utilizing graph representation learning, to effectively and stealthily manipulate federated fine-tuning of LLMs, significantly degrading global model performance while evading standard defenses.

MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection

MIRA proposes a novel source-aware filtering framework that discovers and anchors evaluation rubrics during data selection, significantly improving code-oriented mid-training data quality while reducing token usage.

What Gets Unmasked First? Trajectory Analysis of Diffusion Models for Graph-to-Text Generation

This paper analyzes the decoding process of masked diffusion models for graph-to-text generation, finding that structural fine-tuning disrupts natural entity-first generation and proposing a structural decoding method to fix it.

Efficient Diffusion LLMs via Temporal-Spatial Parallel Decoding and Confidence Extrapolation

The paper proposes a novel trace-aware decoding framework, combining Temporal-Spatial Parallel Decoding (TSPD) and Confidence Extrapolation (CE), to significantly accelerate the inference of diffusion-based LLMs by identifying and fixing converged tokens early.

Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization

The paper proposes a novel render-free framework that conditions video diffusion models directly on compressed 3D human mesh tokens, enabling robust 3D-aware human motion control without relying on rendered 2D guidance.

Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking

The paper introduces Humanoid-GPT, a large-scale generative Transformer model that achieves robust zero-shot motion tracking and control by training on a massive, unified corpus of motion data.

Redesign Mixture-of-Experts Routers with Manifold Power Iteration

This paper proposes a new router redesign for Mixture-of-Experts models using Manifold Power Iteration to align router rows with the principal singular directions of associated experts.

DexCompose: Reusing Dexterous Policies for Multi-Task Manipulation with a Single Hand

A framework called DexCompose is proposed to reuse pretrained dexterous policies for multi-task manipulation with explicit finger-level action ownership.

Chronos: A Physics-Informed Full-History Framework for Non-Markovian Long-Horizon Manipulation

This paper introduces Chronos, a physics-informed framework for non-Markovian long-horizon manipulation, which elevates observation history to the latent state of the policy dynamics and achieves higher success rates and fewer parameters than Markovian VLA baselines in both simulated and real-world experiments.

FedLAB: Traceable Semantic Codebooks for Federated Multimodal Graph Foundation Learning

The paper proposes FedLAB, a traceable semantic codebook framework for federated multimodal graph foundation learning, which organizes multimodal graph knowledge into hierarchical codebooks and refines them through federated semantic barycenter pre-training.

Noisy Environment Adaptation of Neural Speech Codec via Focal Mask and Noise Feature Separation

The paper proposes FocalSE, a method for enhancing speech in neural speech codecs by performing feature denoising, separation, and recognition in the continuous embedding space.

Wat3R: Underwater 3D Geometry Learning without Annotations

This paper proposes Wat3R, a cross-domain semi-supervised learning framework for adapting 3D reconstruction models from air to underwater scenes using unlabeled real underwater video footage and a teacher-student architecture.

DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation

The paper introduces DexVerse, a large-scale and modular benchmark for dexterous manipulation with 100 tasks, 3 robot arms, 6 hands, and configurable visual variations.

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

This paper introduces the REAL-TSE Challenge, a satellite challenge on target speaker extraction from real conversational recordings, and describes its task definition, datasets, evaluation protocol, and submitted systems.

Handroid: Bridging Dexterous Hand and Humanoid

A single robot platform, Handroid, is introduced that can function as both a dexterous hand and a humanoid robot, with interchangeable control and learning frameworks.

Highlighted terms show continued research focus across papers

Papers

cs.ROEmpiricalRecentJul 17, 2026

Handroid: Bridging Dexterous Hand and Humanoid

Ruogu Li, Chenyang Ma, Sikai Li, Zhenyu Wei +5 more

A single robot platform, Handroid, is introduced that can function as both a dexterous hand and a humanoid robot, with interchangeable control and learning frameworks.

View →
eess.AScs.SDEmpiricalRecent
Jul 16, 2026

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

Shuai Wang, Zihan Qian, Ke Zhang, Jiangyu Han +8 more

This paper introduces the REAL-TSE Challenge, a satellite challenge on target speaker extraction from real conversational recordings, and describes its task definition, datasets, evaluation protocol,…

View →
cs.CVEmpiricalRecentJul 9, 2026

Wat3R: Underwater 3D Geometry Learning without Annotations

Jiangwei Ren, Xingyu Jiang, Zijie Song, Wei Xu +3 more

This paper proposes Wat3R, a cross-domain semi-supervised learning framework for adapting 3D reconstruction models from air to underwater scenes using unlabeled real underwater video footage and a tea…

View →
cs.ROEmpiricalRecentJul 9, 2026

DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation

Yunchao Yao, Zhuxiu Xu, Tianqi Zhang, Zixian Liu +11 more

The paper introduces DexVerse, a large-scale and modular benchmark for dexterous manipulation with 100 tasks, 3 robot arms, 6 hands, and configurable visual variations.

View →
eess.ASEmpiricalRecentJul 5, 2026

Noisy Environment Adaptation of Neural Speech Codec via Focal Mask and Noise Feature Separation

Shaokai Li, Weiping Tu, Yuhong Yang

The paper proposes FocalSE, a method for enhancing speech in neural speech codecs by performing feature denoising, separation, and recognition in the continuous embedding space.

View →
cs.LGEmpiricalRecentJun 30, 2026

FedLAB: Traceable Semantic Codebooks for Federated Multimodal Graph Foundation Learning

Zekai Chen, Kairui Yang, Xuaner Chen, Xunkai Li +3 more

The paper proposes FedLAB, a traceable semantic codebook framework for federated multimodal graph foundation learning, which organizes multimodal graph knowledge into hierarchical codebooks and refine…

View →
cs.ROEmpiricalRecentJun 29, 2026

Chronos: A Physics-Informed Full-History Framework for Non-Markovian Long-Horizon Manipulation

Yulin Zhou, Yimeng Wang, Nengyu Wang, Shaojia Xing +8 more

This paper introduces Chronos, a physics-informed framework for non-Markovian long-horizon manipulation, which elevates observation history to the latent state of the policy dynamics and achieves high…

View →
cs.ROcs.AIcs.CVEmpiricalRecentJun 26, 2026

DexCompose: Reusing Dexterous Policies for Multi-Task Manipulation with a Single Hand

Dihong Huang, Zhenyu Wei, Zhuxiu Xu, Yunchao Yao +2 more

A framework called DexCompose is proposed to reuse pretrained dexterous policies for multi-task manipulation with explicit finger-level action ownership.

View →
cs.LGcs.AIcs.CLEmpiricalRecentJun 10, 2026

Redesign Mixture-of-Experts Routers with Manifold Power Iteration

Songhao Wu, Ang Lv, Ruobing Xie, Yankai Lin

This paper proposes a new router redesign for Mixture-of-Experts models using Manifold Power Iteration to align router rows with the principal singular directions of associated experts.

View →
cs.ROcs.AIcs.CVRecentJun 2, 2026

Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking

Zekun Qi, Xuchuan Chen, Dairu Liu, Chenghuai Lin +9 more

The paper introduces Humanoid-GPT, a large-scale generative Transformer model that achieves robust zero-shot motion tracking and control by training on a massive, unified corpus of motion data.

View →
cs.CVcs.AIeess.IVRecentJun 1, 2026

Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization

Jingyun Liang, Min Wei, Shikai Li, Yizeng Han +4 more

The paper proposes a novel render-free framework that conditions video diffusion models directly on compressed 3D human mesh tokens, enabling robust 3D-aware human motion control without relying on re…

View →
cs.CLcs.AIRecentMay 29, 2026

What Gets Unmasked First? Trajectory Analysis of Diffusion Models for Graph-to-Text Generation

Qing Wang, Jacob Devasier, Chengkai Li

This paper analyzes the decoding process of masked diffusion models for graph-to-text generation, finding that structural fine-tuning disrupts natural entity-first generation and proposing a structura…

View →
cs.CLRecentMay 29, 2026

Efficient Diffusion LLMs via Temporal-Spatial Parallel Decoding and Confidence Extrapolation

Zekai Li, Ji Liu, Yiqing Huang, Ziqiong Liu +2 more

The paper proposes a novel trace-aware decoding framework, combining Temporal-Spatial Parallel Decoding (TSPD) and Confidence Extrapolation (CE), to significantly accelerate the inference of diffusion…

View →
cs.AIRecentMay 28, 2026

MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection

Haowen Wang, Yaxin Du, Jian Yang, Jiajun Wu +8 more

MIRA proposes a novel source-aware filtering framework that discovers and anchors evaluation rubrics during data selection, significantly improving code-oriented mid-training data quality while reduci…

View →
cs.LGcs.CRcs.NIRecentMay 8, 2026

Graph Representation Learning Augmented Model Manipulation on Federated Fine-Tuning of LLMs

Hanlin Cai, Kai Li, Houtianfu Wang, Haofan Dong +3 more

The paper proposes an Augmented Model maniPulation (AugMP) strategy, utilizing graph representation learning, to effectively and stealthily manipulate federated fine-tuning of LLMs, significantly degr…

View →
cs.CRcs.SERecentMay 7, 2026

Heimdallr: Characterizing and Detecting LLM-Induced Security Risks in GitHub CI Workflows

Bonan Ruan, Yeqi Fu, Chuqi Zhang, Jiahao Liu +2 more

This paper introduces Heimdallr, a novel framework that characterizes and detects LLM-induced security risks by analyzing the full execution chain of LLM integrations within GitHub CI workflows.

View →
cs.CRcs.AIRecentApr 18, 2026

SafeDream: Safety World Model for Proactive Early Jailbreak Detection

Bo Yan, Weikai Lin, Yada Zhu, Song Wang

SAFEDREAM introduces a lightweight, external world-model framework that proactively detects multi-turn jailbreak attacks by modeling cumulative safety erosion and predicting early failure points.

View →
cs.CRRecentApr 16, 2026

NFTDELTA: Detecting Permission Control Vulnerabilities in NFT Contracts through Multi-View Learning

Hailu Kuang, Xiaoqi Li, Wenkai Li, Zongwei Li

NFTDELTA is a novel framework that uses multi-view learning on static code analysis to detect permission control vulnerabilities in NFT contracts with high accuracy.

View →
cs.CRRecentApr 8, 2026

PSR2: A Phase-based Semantic Reasoning Framework for Atomicity Violation Detection via Contract Refinement

Xiaoqi Li, Xin Wang, Wenkai Li, Zongwei Li

The paper introduces PSR extsuperscript{2}, a novel static analysis framework that significantly improves the detection of atomicity violations in smart contracts by combining structural path searchin…

View →
cs.CRcs.AIcs.CLRecentApr 7, 2026

Say Something Else: Rethinking Contextual Privacy as Information Sufficiency

Yunze Xiao, Wenkai Li, Xiaoyuan Wu, Ningshan Ma +2 more

The paper proposes Information Sufficiency (IS) as a comprehensive framework for privacy-preserving LLM communication, demonstrating that free-text pseudonymization outperforms existing suppression an…

View →