pentesting is an autonomous security agent built for offensive security learning, CTF competitions, and real-world penetration testing workflows.
Recorded autonomous attempts; scores select the newest finalized attempt per task.
| Model | Solved | Rate | Tokens | Avg Time | Est. Cost | L1 / L2 / L3 Solved |
|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash | 102 / 104 | 98.1% | 257.0M | 19.6 min | Unknown | 44/45 ยท 51/51 ยท 7/8 |
| GLM-5.3-Flash | 94 / 104 | 90.4% | 176.6M | 22.6 min | Unknown | 44/45 ยท 44/51 ยท 6/8 |
Scores are not single-attempt success rates. Tokens and average time include all retained finalized attempts (DeepSeek: 217; GLM: 214); time is summed attempt duration divided by 104, not campaign wall time. Complete cost is unavailable because prices or attempt measurements are missing. See the DeepSeek report and GLM report for selection and measurement limits.
# .env
OPENAI_API_KEY="your-api-key"
OPENAI_MODEL="your-model-name"
OPENAI_BASE_URL="https://api.openai.com/v1" # optional, e.g. https://openrouter.ai/api/v1
OPENAI_CONTEXT_TOKENS="128k" # optional context ceiling, e.g. 128k, 1m
OPENAI_MAX_OUTPUT_TOKENS="16k" # optional max output per turn, e.g. 16k, 32k
# Option A: Docker Compose (Recommended)
docker compose --env-file .env -f docker/compose.yaml run --rm pentesting
# Option B: Docker CLI โ Interactive TUI
docker run --rm -it --init \
--cap-add=NET_RAW --cap-add=NET_ADMIN \
--env OPENAI_API_KEY="your-api-key" \
--env OPENAI_MODEL="your-model-name" \
--env OPENAI_CONTEXT_TOKENS="128k" \
--env OPENAI_MAX_OUTPUT_TOKENS="16k" \
-v ${PWD}/workspace:/workspace \
-v ${PWD}/runs:/state \
agnusdei1207/pentesting:latest
# Option C: Docker CLI โ One-shot Goal
docker run --rm -it --init \
--cap-add=NET_RAW --cap-add=NET_ADMIN \
--env OPENAI_API_KEY="your-api-key" \
--env OPENAI_MODEL="your-model-name" \
--env OPENAI_CONTEXT_TOKENS="128k" \
--env OPENAI_MAX_OUTPUT_TOKENS="16k" \
-v ${PWD}/workspace:/workspace \
-v ${PWD}/runs:/state \
agnusdei1207/pentesting:latest \
run --goal "Investigate the target and solve the objective" --workspace /workspace --run /state/current
The v0.200.8 native release provides Linux x64 binaries (glibc 2.36 or newer). Use the Docker image on other platforms; native Windows, macOS, and ARM64 assets are not included in this release.
# Option A: Build and launch isolated Docker TUI
npm run check
# Option B: Global CLI โ Interactive TUI
npm install --global pentesting
pentesting
# Option C: Global CLI โ One-shot Goal
pentesting run --goal "Investigate the target and solve the objective" --workspace .
Anyone interested in offensive security and autonomous agent engineering is warmly welcome. From simple typo fixes and domain tradecraft insights to bug reports and pull requests, every contribution is appreciated.
Feel free to open an issue or submit a pull request. We look forward to building a sharper, more reliable tool together.
powershell -File scripts/nverify.ps1.MIT