This paper proposes a four-layer technical architecture for large model inference optimization, including Multi-model Fusion, Model Optimization, Compute-Model Fusion, and Compute-Network-Model Fusion.
Proposes a new technical architecture for large model inference optimization
Keywords
Before reading this…
Applications
To understand this paper, make sure you know these concepts first:
Large model inference optimization serves as a key foundation for supporting the scalable, low-cost, and highly stable operation of large model services. Centered on token-oriented inference optimization technology, this paper proposes for the first time a four-layer technical architecture consisting of Multi-model Fusion, Model Optimization, Compute-Model Fusion, and Compute-Network-Model Fusion. It systematically reviews the key technologies and current industry status across these four levels and analyzes the application value of related technologies in real-world business scenarios. This paper provides a practical technical path for reducing token production costs, improving token service efficiency, ensuring the stability of token supply, and driving the transition of large model services from being merely callable to being operable.
n-VM: A Multi-VM Layer-1 Architecture with Shared Identity and Token State
The paper proposes n-VM, a novel Layer-1 architecture that unifies multiple hete…
CompRank: Efficient LLM Reranking via Token-Level Compression and Decoding-Free Scoring
This paper proposes CompRank, a token-efficient reranking framework for large la…
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
The paper proposes TriageFuzz, a token-aware fuzzing framework that significantl…
Privacy Guard & Token Parsimony by Prompt and Context Handling and LLM Routing
The paper introduces a 'Privacy Guard' framework that simultaneously reduces ope…
GUARD-SLM: Token Activation-Based Defense Against Jailbreak Attacks for Small Language Models
The paper proposes GUARD-SLM, a token activation-based defense mechanism, to enh…
NANOZK: Layerwise Zero-Knowledge Proofs for Verifiable Large Language Model Inference
NANOZK introduces a novel, highly efficient zero-knowledge proof system that all…
AUTOGATE: Automated Clock Gating via Toggling-Aware LLM-based RTL Rewriting
This paper introduces AUTOGATE, an agentic framework for industry-grade RTL powe…
DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?
This paper introduces DIRECT, a routing framework that allocates test-time compu…