* fix(auth0): harden against security-scan false positives The v2.0 re-architecture (#137) triggered a critical finding on the Agent Trust Hub / Socket / Snyk scans. This addresses the actionable, scan-scoped surface without removing legitimate behavior. Placeholder URLs - Replace registrable placeholder hostnames (your-api.com, your-production-domain.com, your-app.com, your-spa-domain.com, your-web-app.com, your-domain.com) with RFC-2606 example.com forms across references/*.md. Kills the "malicious URL" finding (your-production-domain.com) and pre-empts future URL-reputation hits. Matches the skill's already-dominant `api.example.com` convention. .auth0.com domain placeholders and client-id/secret tokens are left as-is: not registrable / not URLs, so not a scanner surface. Move the behavioral eval harness out of the skill dir - Socket flagged tests/behavioral/graders.mjs (reads .env, pipes source into the local `claude` CLI via execa). That's a dev-only test grader, but it sits inside per-skill scan scope. Move tests/ -> repo-root evals/ so the executable harness is no longer bundled with the skill consumers install. Fix run-evals.mjs path resolution, the routing checker's routing-cases.json lookup, the README, and the AGENTS.md layout doc. skillsaw: exempt template-routed reference families - framework-*.md / tooling-*.md are routed only via the `references/framework-{framework}.md` / `tooling-{tooling}.md` placeholder templates, which skillsaw's literal-scan agentskill-unreferenced-files rule can't resolve. Co-located tests were previously (accidentally) satisfying that reachability; moving them out exposed it. Their reachability is properly enforced by scripts/check_router_reachability.py, so exempt the two template-routed families in .skillsaw.yaml. feature-*/pattern-* orphan detection stays. Verified: skillsaw --strict (0/0, grade A), check_router_reachability, check_routing_evals, and run-evals --dry-run (17 case files) all pass. * fix(auth0): keep load_cases skill-dir-relative for unit tests The previous commit hardcoded routing-cases.json to repo-root evals/, which broke scripts/test_check_routing_evals.py — the unit tests build a self-contained temp skill with its own tests/routing-cases.json and rely on load_cases(skill_dir) reading from that dir. Prefer a skill-local tests/routing-cases.json when present (unit tests), fall back to repo-root evals/routing-cases.json otherwise (production, after the harness move). Both the pytest suite (10/10) and the real check pass. * chore(auth0): bump plugin + skill version to 2.0.1 Patch bump for the security-scan hardening: placeholder-URL cleanup, eval harness move out of the skill dir, and the skillsaw template-route exemption. No new routes or breaking changes. Updates all six plugin/ marketplace manifests and SKILL.md frontmatter in lockstep.
Behavioral evals — unified auth0 skill
Two layers of eval guard this skill. This directory is the behavioral layer;
the routing layer lives in ../routing-cases.json + scripts/check_routing_evals.py.
| Layer | Question | Runs in CI? | Needs a live model? |
|---|---|---|---|
Routing (../routing-cases.json) |
Does each intent/framework map to reference files that exist? | ✅ yes | no — deterministic |
| Behavioral (here) | Does the skill make the agent generate correct SDK code? | ❌ manual | yes — claude CLI |
Behavioral evals drive a real agent and grade the code it writes, so they need
the claude CLI and (for cases that configure a tenant) real Auth0 credentials
in the prompt. That's why they're a manually-run suite, not CI.
History
These consolidate the per-skill tests/ harnesses that shipped with the ~44
individual skills before the single-skill migration. Each skill carried its own
near-identical ~1,200-line run-evals.mjs; that logic now lives once in
graders.mjs + run-evals.mjs, with one case file per framework/feature under
cases/. Each case records the origin_skill it came from.
Layout
behavioral/
├── run-evals.mjs # the single runner (drives the agent, reports deltas)
├── graders.mjs # shared grader engine (contains/matches/judge/...)
├── package.json # execa dependency
└── cases/
├── flask.json # { slug, origin_skill, evals[], graders[] }
├── express-jwt.json
└── ... # 17 cases; 13 have machine graders,
# 4 (branding, custom-domains, cli, acul) are
# expectations-only → manual transcript review
Running
cd evals/behavioral
npm install # once, for execa
node run-evals.mjs --list # show cases
node run-evals.mjs --dry-run # validate case files + grader regexes (no agent)
node run-evals.mjs # run every case
node run-evals.mjs flask express-jwt # run only named slugs
node run-evals.mjs --model <id> # pin a model for the agent AND judge graders
node run-evals.mjs --skill-only # skip the without-skill comparison
For each graded case the runner runs the prompt with the skill (loads the
auth0 plugin via --plugin-dir, so the router activates and detects the
framework) and without it, grades both workspaces, and prints the delta — a
useful skill should score materially higher with the router loaded.
Adding / updating a case
Edit the JSON under cases/. Shape:
{
"slug": "flask", // matches the router's framework/feature slug
"origin_skill": "auth0-flask",
"evals": [{ "prompt": "...", "expectations": ["..."] }],
"graders": [ // omit (or null) for a manual-review case
{ "type": "matches", "pattern": "ServerClient", "description": "SDK initialized" },
{ "type": "not_contains_any", "values": ["Authlib", "python-jose"], "description": "no wrong lib" },
{ "type": "judge", "question": "Does the code ...?", "examples": "PASS: ...\nFAIL: ..." }
],
"scaffold": { // optional — seed files so Tier 1/2 detection fires
"package.json": "{ \"dependencies\": { \"@auth0/nextjs-auth0\": \"^4.0.0\" } }"
}
}
Grader types: contains, contains_any, not_contains, not_contains_any,
matches (regex), not_matches (regex must be absent), file_contains
(file_pattern glob + value), all (composite), judge (LLM YES/NO).
Notes on the negative graders:
- Prefer
not_matchesovernot_containswhen a bare substring would false-positive. E.g. asserting the deprecatedexpress-jwtpackage is absent: a plainnot_containsfor"express-jwt"also fires on the correct packageexpress-oauth2-jwt-bearer(via its dep graph) and on natural project names likeexpress-jwt-api. A regex like["']express-jwt["'](dep key / import target only) avoids that. - Generated lockfiles (
package-lock.json,Podfile.lock,*.lock, …) are excluded from the source scan entirely — they pin the full transitive graph, so substrings there don't reflect the authored code. not_contains*/not_matchesgraders are auto-invalidated if no positive grader passed, so an empty workspace can't score by writing nothing.
The judge grader asks a live model for a VERDICT: YES/NO (parsed from the
end of the reply, so a judge that reasons before concluding is read correctly).
Don't hardcode a specific SDK version in a grader — the references deliberately
teach "use the current version," so a "^1.7.4"-style pin tests removed advice
and rots on every release. Match "a version is present" only if you must.