Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence Pool | ArxivCSExplorer