SkillTester: Benchmarking Utility and Security of Agent Skills
SkillTester is a comprehensive tool and framework designed to benchmark both the functional utility and the security robustness of agent skills, providing standardized scores and status labels.
Abstract
More Like ThisThis technical report presents SkillTester, a tool for evaluating the utility and security of agent skills. Its evaluation framework combines paired baseline and with-skill execution conditions with a separate security probe suite. Grounded in a comparative utility principle and a user-facing simplicity principle, the framework normalizes raw execution artifacts into a utility score, a security score, and a three-level security status label. More broadly, it can be understood as a comparative quality-assurance harness for agent skills in an agent-first world. The public service is deployed at https://skilltester.ai, and the broader project is maintained at https://github.com/skilltester-ai/skilltester.