Back to Paper
cs.CR
Local ID: 2604.27601v2
AI Summary: gemma4:e4b
SecGoal: A Benchmark for Extracting Formalizable Security Goals from Protocol Documents
By Dawei Huang, Hui Li, Bo Jia, Haonan Feng, Jingjing Guan, Yueshuang Jiao, Xiangdong Li
Revision History Timeline
v14/30/2026
4/30/2026
“19 pages, 4 figures”
v25/28/2026
5/28/2026
“23 pages, 4 figures”
★ Version indexed in ExplorerComparing v1 vs v2
Green = Added • Red = Removed
Title Comparison
SecGoal: A Benchmark for Security GoalExtracting ExtractionFormalizable andSecurity FormalizationGoals from Protocol Documents
Authors Comparison
Added:Xiangdong Li
Unchanged:Dawei Huang, Hui Li, Bo Jia, Haonan Feng, Jingjing Guan, Yueshuang Jiao
v1 Comment
“19 pages, 4 figures”
v2 Comment
“23 pages, 4 figures”
Abstract Word Diff
Formal verification provides rigorous guarantees for cryptographic security, yet automating the extraction and formalizationextracting offormalizable security goals from natural languagenatural-language protocol documents remains a majorlargely bottleneck,manual. compoundedWe byintroduce theSecGoal, scarcitya ofdedicated expert-annotated resourcesdataset and integrated frameworks bridging unstructured text andbenchmark symbolicfor logic.extracting Weformalizable introducesecurity SecGoal,goal thestatements firstfrom expert-annotatedprotocol benchmarkdocuments, covering 15 widely deployed protocolprotocols, documents,together includingwith 5G-AKAAIFG, a schema- and TLSflow-conditioned 1.3,framework andfor AIFG,structured anformal AI-assistedsecurity frameworkproperty generation. Our evaluation shows that decomposesfrontier theand tasklarge intoLLMs context-awareachieve goalhigh extractionproperty andrecall retrieval-augmentedbut formalization.low Weextraction conductprecision abecause comprehensivethey evaluationoften fail to assessdistinguish whetherformalizable contemporarysecurity LLMsgoals arefrom readynon-goal toprotocol automatecontent. thisIn pipeline.contrast, OurSecGoal resultsfine-tuning revealmakes asmaller pronouncedopen-source precision-recallLLMs imbalance:substantially frontiermore models,selective suchextractors asof Geminiformalizable 2.5-Pro,security achievegoals. highOn recallthe butheld-out test protocols, Gemma2-9B-FT improves extraction precision belowfrom 15%,24.0\% frequentlyto misclassifying66.6\% operationaland textreaches as97.6\% securityproperty goals.recall, outperforming larger prompted LLMs and encoder baselines. In contrast,a instructioncontrolled tuningsetting, onAIFG SecGoalshows enablesthat compactconcise modelsgoal withinputs 7B/9Bcan parameterssupport tohigh-recall achievestructured F1-scoresproperty abovegeneration, 80%,while substantiallyexpert-vetted outperformingextracted largerinputs general-purposereveal models.over-generation Ouras workthe establishesmain aremaining foundationalbottleneck. datasetTogether, SecGoal and reproducibleAIFG baselineprovide a dataset, benchmark, and framework for automatedspecification-grounded formalsecurity protocolgoal analysis.extraction and property generation.