Files
Frederik Prijck 0d42696025 fix(auth0): harden against security-scan false positives (#144)
* fix(auth0): harden against security-scan false positives

The v2.0 re-architecture (#137) triggered a critical finding on the
Agent Trust Hub / Socket / Snyk scans. This addresses the actionable,
scan-scoped surface without removing legitimate behavior.

Placeholder URLs
- Replace registrable placeholder hostnames (your-api.com,
  your-production-domain.com, your-app.com, your-spa-domain.com,
  your-web-app.com, your-domain.com) with RFC-2606 example.com forms
  across references/*.md. Kills the "malicious URL" finding
  (your-production-domain.com) and pre-empts future URL-reputation hits.
  Matches the skill's already-dominant `api.example.com` convention.
  .auth0.com domain placeholders and client-id/secret tokens are left
  as-is: not registrable / not URLs, so not a scanner surface.

Move the behavioral eval harness out of the skill dir
- Socket flagged tests/behavioral/graders.mjs (reads .env, pipes source
  into the local `claude` CLI via execa). That's a dev-only test grader,
  but it sits inside per-skill scan scope. Move tests/ -> repo-root
  evals/ so the executable harness is no longer bundled with the skill
  consumers install. Fix run-evals.mjs path resolution, the routing
  checker's routing-cases.json lookup, the README, and the AGENTS.md
  layout doc.

skillsaw: exempt template-routed reference families
- framework-*.md / tooling-*.md are routed only via the
  `references/framework-{framework}.md` / `tooling-{tooling}.md`
  placeholder templates, which skillsaw's literal-scan
  agentskill-unreferenced-files rule can't resolve. Co-located tests
  were previously (accidentally) satisfying that reachability; moving
  them out exposed it. Their reachability is properly enforced by
  scripts/check_router_reachability.py, so exempt the two template-routed
  families in .skillsaw.yaml. feature-*/pattern-* orphan detection stays.

Verified: skillsaw --strict (0/0, grade A), check_router_reachability,
check_routing_evals, and run-evals --dry-run (17 case files) all pass.

* fix(auth0): keep load_cases skill-dir-relative for unit tests

The previous commit hardcoded routing-cases.json to repo-root evals/,
which broke scripts/test_check_routing_evals.py — the unit tests build a
self-contained temp skill with its own tests/routing-cases.json and rely
on load_cases(skill_dir) reading from that dir.

Prefer a skill-local tests/routing-cases.json when present (unit tests),
fall back to repo-root evals/routing-cases.json otherwise (production,
after the harness move). Both the pytest suite (10/10) and the real
check pass.

* chore(auth0): bump plugin + skill version to 2.0.1

Patch bump for the security-scan hardening: placeholder-URL cleanup,
eval harness move out of the skill dir, and the skillsaw template-route
exemption. No new routes or breaking changes. Updates all six plugin/
marketplace manifests and SKILL.md frontmatter in lockstep.
2026-07-17 14:14:15 +02:00
..

Behavioral evals — unified auth0 skill

Two layers of eval guard this skill. This directory is the behavioral layer; the routing layer lives in ../routing-cases.json + scripts/check_routing_evals.py.

Layer Question Runs in CI? Needs a live model?
Routing (../routing-cases.json) Does each intent/framework map to reference files that exist? ✅ yes no — deterministic
Behavioral (here) Does the skill make the agent generate correct SDK code? ❌ manual yes — claude CLI

Behavioral evals drive a real agent and grade the code it writes, so they need the claude CLI and (for cases that configure a tenant) real Auth0 credentials in the prompt. That's why they're a manually-run suite, not CI.

History

These consolidate the per-skill tests/ harnesses that shipped with the ~44 individual skills before the single-skill migration. Each skill carried its own near-identical ~1,200-line run-evals.mjs; that logic now lives once in graders.mjs + run-evals.mjs, with one case file per framework/feature under cases/. Each case records the origin_skill it came from.

Layout

behavioral/
├── run-evals.mjs     # the single runner (drives the agent, reports deltas)
├── graders.mjs       # shared grader engine (contains/matches/judge/...)
├── package.json      # execa dependency
└── cases/
    ├── flask.json        # { slug, origin_skill, evals[], graders[] }
    ├── express-jwt.json
    └── ...               # 17 cases; 13 have machine graders,
                          # 4 (branding, custom-domains, cli, acul) are
                          # expectations-only → manual transcript review

Running

cd evals/behavioral
npm install            # once, for execa

node run-evals.mjs --list          # show cases
node run-evals.mjs --dry-run       # validate case files + grader regexes (no agent)
node run-evals.mjs                 # run every case
node run-evals.mjs flask express-jwt   # run only named slugs
node run-evals.mjs --model <id>    # pin a model for the agent AND judge graders
node run-evals.mjs --skill-only    # skip the without-skill comparison

For each graded case the runner runs the prompt with the skill (loads the auth0 plugin via --plugin-dir, so the router activates and detects the framework) and without it, grades both workspaces, and prints the delta — a useful skill should score materially higher with the router loaded.

Adding / updating a case

Edit the JSON under cases/. Shape:

{
  "slug": "flask",              // matches the router's framework/feature slug
  "origin_skill": "auth0-flask",
  "evals": [{ "prompt": "...", "expectations": ["..."] }],
  "graders": [                  // omit (or null) for a manual-review case
    { "type": "matches", "pattern": "ServerClient", "description": "SDK initialized" },
    { "type": "not_contains_any", "values": ["Authlib", "python-jose"], "description": "no wrong lib" },
    { "type": "judge", "question": "Does the code ...?", "examples": "PASS: ...\nFAIL: ..." }
  ],
  "scaffold": {                 // optional — seed files so Tier 1/2 detection fires
    "package.json": "{ \"dependencies\": { \"@auth0/nextjs-auth0\": \"^4.0.0\" } }"
  }
}

Grader types: contains, contains_any, not_contains, not_contains_any, matches (regex), not_matches (regex must be absent), file_contains (file_pattern glob + value), all (composite), judge (LLM YES/NO).

Notes on the negative graders:

  • Prefer not_matches over not_contains when a bare substring would false-positive. E.g. asserting the deprecated express-jwt package is absent: a plain not_contains for "express-jwt" also fires on the correct package express-oauth2-jwt-bearer (via its dep graph) and on natural project names like express-jwt-api. A regex like ["']express-jwt["'] (dep key / import target only) avoids that.
  • Generated lockfiles (package-lock.json, Podfile.lock, *.lock, …) are excluded from the source scan entirely — they pin the full transitive graph, so substrings there don't reflect the authored code.
  • not_contains* / not_matches graders are auto-invalidated if no positive grader passed, so an empty workspace can't score by writing nothing.

The judge grader asks a live model for a VERDICT: YES/NO (parsed from the end of the reply, so a judge that reasons before concluding is read correctly).

Don't hardcode a specific SDK version in a grader — the references deliberately teach "use the current version," so a "^1.7.4"-style pin tests removed advice and rots on every release. Match "a version is present" only if you must.