20 results for “Desktop-Delta Bench (DDB)”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper introduces Desktop-Delta Bench (DDB), an offline step-level benchmark for evaluating computer-use agents' ability to reconstruct causal transitions in desktop GUI environments.
This paper introduces DGNA, a methodology to unveil the Non-Uniform Memory Access (NUMA) architecture of GPU memory hierarchy through microbenchmarking and data analysis.
Junming Chen, Junyang Jiang, Xu Chen, Zibo Liang +1 more
The paper introduces DBA-Bench, a benchmark for evaluating database agents with production fidelity, outcome-first evaluation, and controlled scenario reproducibility.
The paper introduces FinVerBench, a comprehensive benchmark for financial statement verification, concluding that successful verification requires calibrated judgment under realistic observational con…
Minglai Yang, Xinyan Velocity Yu, Pengyuan Li, Xinyu Guo +21 more
The paper introduces Dr. DocBench, a difficulty-aware, comprehensive benchmark designed to rigorously test expert-level and challenging document parsing capabilities for VLMs, demonstrating that curre…
RefDiffNet is a lightweight, plug-and-play module that enhances PCB defect detection by comparing the defective image to a defect-free reference image, significantly improving detection accuracy with…
Haoyuan Tang, Zhuo Zhang, Jialin Li, Shuai Xiao +1 more
The paper proposes VDSB-GWSyn, a Diffusion Schrödinger Bridge framework, to synthesize controllable and anatomically feasible guidewire images on coronary angiography (CAG) scans, significantly improv…
The paper introduces SpatialBench-Long, a comprehensive benchmark designed to test AI agents' ability to perform end-to-end scientific reasoning and derive biological claims from complex, raw spatial…
The paper introduces PortBench, a comprehensive benchmark that evaluates LLMs for portfolio management by assessing both correlation awareness and performance across a full, multi-stage decision pipel…
Dezhi Yi, Huifeng Guo, Kunpeng Xie, Zhaolong Jian +5 more
This paper proposes DPIFrame, a dual parallelizable framework for accelerating Click-through rate (CTR) model inference on GPU, achieving state-of-the-art inference performance with significant speedu…
The paper introduces MUSE, a comprehensive benchmark that evaluates Text-to-CAD generation by assessing complex assemblies based on functionality, manufacturability, and assemblability, moving beyond…
This paper provides the first large-scale characterisation of Domain-Driven Design (DDD) adoption and implementation on GitHub.
Xiang Wang, Tingting Zhang, Sen Wang, Ying Wu +3 more
The paper introduces PetroBench, a comprehensive benchmark for evaluating Large Language Models across various domains of petroleum engineering, finding that models perform better on subjective tasks…
The paper proposes a Doeblin-anchored contrastive chart to learn valid Markov transition kernels by combining the target transition with a restart law, ensuring the learned object is mathematically so…
The paper introduces CaDDTree, a cost-aware method that optimizes token throughput by jointly selecting the tree structure and node budget for speculative decoding, outperforming existing methods like…
The paper introduces JupOtter, a bug detection system for Jupyter Notebooks with three contributions: notebook-specific tokenization, cell-level bug prediction, and a labeled dataset called OtterDatas…
The paper demonstrates that for FFT-based radar imaging on Apple Silicon, the limiting factor for half-precision (FP16) is dynamic range, not mantissa precision, and proposes a block-floating-point (B…