My
50 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
CORTIS is a text-only adaptation framework that fine-tunes spoken language models for task-oriented voice agents using text-form task supervision.
This paper proposes DysLexLens, a framework to analyze dyslexic learners' experiences with AI using a low-resource LLM, with features including dictionary-driven filtering, semantic analysis, quantitative evaluation metrics, and qualitative validation guidelines.
The paper conducts a reproducibility study on FACTER, a model-agnostic framework for fairness and statistical coverage in LLM-based recommendation, and evaluates its consistency and contribution.
This paper analyzes short- to medium-term HDD failure rates of HGST, Seagate, Toshiba, and Western Digital using the Backblaze dataset.
This paper introduces TRACE, a model for interpretable 4-class glioblastoma response classification on longitudinal 3D MRI using a structured concept reasoning approach.
This paper evaluates the improvement of statistical fault localization by augmenting it with execution features.
The paper introduces DigitalCoach, a dataset of human expert-novice computer use coaching sessions, and evaluates the ability of state-of-the-art models to teach humans how to use computers.
This paper introduces MAGNET, a framework for long-form narrative generation and verification using a multi-agent goal-driven engine and a graph-based pipeline.
The paper constructs datasets and benchmarks to evaluate the performance of visual generators in handling open-ended requests, and proposes a teach-then-search co-training framework to improve their world-knowledge.
The paper introduces Graph-as-Policy (GaP), a multi-agent coding harness for Variational Automation tasks that generates directed computation graphs and improves success rates and throughput through internal simulation.
An industrial evaluation of an AI-enabled decision-support tool for crime linkage analysis was conducted, showing analysts selectively used AI predictions and valued model features.
This paper introduces Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scales visuomotor context to 8K timesteps, enabling new capabilities like one-shot imitation and robustness to perturbations.
This paper proposes a thermodynamic computing stack for machine learning using stochastic analog processes and energy-based models.
This paper introduces Retriever, an asynchronous decision model and runtime system for building long-horizon robot agents with explicit clock and input-consumption semantics.
RAMP is a method for improving click-through rate and conversion rate prediction accuracy in privacy-constrained settings by using a personalized pathway, a non-personalized pathway, and a prediction-alignment architecture.
This paper proposes a lightweight, distributed machine learning framework for sub-centimeter indoor localization using D-MIMO in O-RAN architectures, reducing midhaul traffic by 100x while maintaining accuracy.
This paper proposes a system called Autogram that uses AI and statistics to discover and validate network invariants.
This paper proposes a physics-guided framework, PathRIR, for fast room impulse response simulation using image-source-method, preserving geometric structure while pruning acoustically insignificant paths and compensating for energy loss.
This paper proposes a proactive fair scheduling framework to prevent fabric oversubscription and eliminate incast in Mixture of Experts (MoE) architectures, demonstrating consistent link utilization and reduced Collective Completion Time (CCT) through simulations.
This paper presents NELSSA, a serving system that integrates GPUs with Processing-near-Memory (PNM) devices to efficiently handle mixed-length workloads in LLMs, achieving up to 5.5x decode throughput improvement and 15x P99 latency reduction.
Papers
NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement
Sookyung Choi, Seungyong Lee, Kangkyu Park, Yunseo Chun +10 more
This paper presents NELSSA, a serving system that integrates GPUs with Processing-near-Memory (PNM) devices to efficiently handle mixed-length workloads in LLMs, achieving up to 5.5x decode throughput…