20 results for “long thetas”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper investigates limitations of learning tanh neural networks under finite-precision computations and Lp accuracy guarantees.
This paper introduces the concept of 'large' zeros in linear recurrence sequences and establishes that they either do not exist or are very sparse, which could lead to decidability of the Skolem Probl…
CaMBRAIN introduces a novel Mamba-based State Space Model (SSM) for real-time, continuous EEG inference, achieving state-of-the-art results with significantly higher throughput than existing methods.
The paper proposes Periodic RoPE (P-RoPE) combined with a dual-layer attention mechanism to overcome the positional encoding limitations of LLMs, enabling theoretically infinite context understanding.
Chen He, Yuhao Wu, Lei Wang, Wenxuan Zhang +1 more
The paper identifies and demonstrates that post-conclusion continuation in answer-correct long-CoT traces is harmful during LLM fine-tuning, proposing a method to cut this continuation.
HARP introduces a novel, adaptive, learnable orthogonal processor that significantly improves the robustness and accuracy of extreme low-bit LLM quantization compared to fixed methods.
LongAttnComp introduces a novel, two-stage fine-tuning framework for context compression that significantly improves long-context reasoning performance, matching or exceeding full-context accuracy on…
This paper presents an efficient algorithm for right-to-left sequence prediction based on a new complexity measure called arithmetic repetition complexity, and demonstrates its application to predicti…
The paper analyzes the structured CVP distance on the log-unit lattice of cyclotomic fields, significantly reducing the conjectured CDPR factor for the ML-KEM cryptosystem from exponential to sub-poly…
Tieyao Zhang, Yuke Liu, Jiaxing Yu, Xinda Wu +2 more
This paper proposes RPPNet, a two-stage deep learning architecture for music generation with variable structural boundaries, which automatically derives grouping of Rhythm-Pitch Primitive sequences fr…
The paper introduces a modular version of rank and linear-complexity tests for pseudorandom number generators and provides a Rust program, modlin, to detect statistical bias in generators that are lin…
This paper explains how Rotary Position Embeddings (RoPE) frequencies in transformer models correspond to the relative-distance structure of training data, and the implications for long-context genera…
The paper introduces BRo-JEPA, a latent world model that successfully learns modular arithmetic (like addition modulo 10) by explicitly imposing the circular structure of the problem into the latent s…
SHARP proposes a novel sleep-based hierarchical replay framework to efficiently learn long-range non-stationary temporal patterns in streaming data, achieving improved context retention and predictive…
The paper introduces Morlet Positional Encoding (MoPE), a novel wavelet-based positional encoding that models position and locality simultaneously, outperforming standard sinusoidal and RoPE methods.
This paper disproves a conjecture about the longest-path lengths of MaxSTN and MinSTN algorithms for producing $st$-orientations of biconnected graphs.
This paper presents two algorithms using the Newton-Raphson method to calculate the floor of y**(1/m) for natural integer numbers y > 2 and m > 1, which can be used to determine if y is an integer pow…
The paper analyzes a fragment of Higher-Order Datalog, showing that restricting recursion to a linear form shifts its expressive power from time complexity to space complexity, specifically capturing…
The paper proposes a novel set of combined cellular automaton (CA)-based pseudo-random number generators (PRNGs) that overcome the weak equidistribution issues of existing CA-based PRNGs, achieving ma…