mirror of
https://github.com/CopilotKit/CopilotKit.git
synced 2026-09-14 16:26:20 +08:00
84dd86f2ed
The marked block that wires managed Intelligence is the region a hosted reader copies verbatim, and nothing checked it. Both gaps were deliberate: the parity manifest lists `src/app/api/copilotkit/**` under `allowedDivergence` for every instance it tracks, and no `docker-compose.test.yml` sets `COPILOTKIT_LICENSE_TOKEN`, so every smoke-tested starter takes the else arm and the `intelligence:` arm has never run in CI. The cost was already visible. The block's code was byte-identical in 21 of 22 starters, but its warning comment had drifted into five variants and the two `ms-agent-framework-*` starters shipped the `demo-user` stub with no warning at all. That drift is how the localhost default of OSS-981 survived in all 22 copies at once. Add `scripts/validate-intelligence-wiring-block.ts`, which greps the opening marker, compares every site against the north-star starter, and fails on the first line that differs. Two normalisations keep it usable: the block is dedented, because `agentcore` nests it deeper, and the else arm's runner name is masked, because `agentcore` runs `AgentCoreRunner` in front of a Bedrock session where an in-process runner has nothing to run. Everything else, comment text included, must match to the byte. Then unify the warning at all 22 sites on the fullest wording, which also says the id must exist in Intelligence or thread operations can fail. The check passes on day one, so it is a ratchet rather than a migration. It is a shape gate, not a content gate: 22 identically wrong copies still pass. What it guarantees is that a fix reaches all of them or none. Not covered: enrolling the `intelligence:` arm in the smoke path. That needs a license token in CI and a reachable endpoint from the compose network, and is tracked separately.
140 lines
5.4 KiB
Markdown
140 lines
5.4 KiB
Markdown
# `_parity/` — integration-demo parity tooling
|
|
|
|
Keeps `examples/integrations/*` demos aligned to a single north-star so
|
|
drift doesn't pile up as the canonical demo evolves.
|
|
|
|
**North-star (v1):** `langgraph-python` — the richest demo (todos state, 5
|
|
tools, polished prompt, full frontend canvas). Every other integration demo
|
|
should track it.
|
|
|
|
## What gets tracked
|
|
|
|
Declared in [`manifest.json`](./manifest.json):
|
|
|
|
- **verbatim files** — copied byte-for-byte from north-star to each instance
|
|
(frontend components, hooks, lib, public assets, shared Docker files,
|
|
`docker-compose.test.yml`, `entrypoint.sh`, `postcss.config.mjs`, etc.)
|
|
- **package.json keys** — tracked dependency versions and script names
|
|
(`@copilotkit/*`, `next`, `react`, shared dev scripts). Per-instance
|
|
overrides in manifest (`packageJsonOverrides`) win where the instance
|
|
legitimately differs (e.g. `dev:agent` runs `npm install` in JS but
|
|
`uv sync` in FastAPI).
|
|
- **canonical prompt** — `_parity/canonical/PROMPT.md`. Each agent
|
|
**inlines** the prompt as a string literal in its source (matching the
|
|
north-star's `main.py` pattern). Verifier greps the canonical prompt's
|
|
first non-blank line against instance agent source — drift = error.
|
|
- **agent surface** — tool names + state keys expected to appear in each
|
|
instance's agent source. Grep-level check — doesn't validate call-site
|
|
correctness, that's the aimock fixture tests' job.
|
|
|
|
## What doesn't get tracked (allowed divergence)
|
|
|
|
Per-instance `allowedDivergence` list in `manifest.json`:
|
|
|
|
- `agent/**` — agents are written in different languages/runtimes (Python
|
|
create_agent, TS StateGraph, Python StateGraph+FastAPI). Human-authored.
|
|
- `src/app/api/copilotkit/**` — north-star uses `LangGraphAgent`, Docker
|
|
instances use `HttpAgent`. Different routes. The Intelligence wiring block
|
|
inside those routes is held to one shape by
|
|
`scripts/validate-intelligence-wiring-block.ts`, which covers all 22 starters
|
|
rather than the 7 tracked here (OSS-982).
|
|
- `Dockerfile`, `docker/Dockerfile.agent`, `serve.py`, `scripts/**` —
|
|
language-specific build/run tooling.
|
|
|
|
Anything outside both `tracked` and `allowedDivergence` is "no-op" — the
|
|
verifier neither checks nor touches it.
|
|
|
|
## Commands
|
|
|
|
From the repo root:
|
|
|
|
```bash
|
|
# Sync a single instance to north-star (copies verbatim files, rewrites
|
|
# package.json keys, writes canonical prompt to agent/PROMPT.md)
|
|
pnpm parity:sync --target=langgraph-js
|
|
|
|
# Dry-run: show what would change without writing
|
|
pnpm parity:sync --target=langgraph-js --dry-run
|
|
|
|
# Sync every non-north-star instance
|
|
pnpm parity:sync --all
|
|
|
|
# Verify — exits non-zero on unexpected drift
|
|
pnpm parity:verify
|
|
pnpm parity:verify --target=langgraph-js
|
|
|
|
# CI invocation (same as verify, no color)
|
|
pnpm parity:check
|
|
```
|
|
|
|
## Typical workflows
|
|
|
|
### North-star changed — sync instances
|
|
|
|
```bash
|
|
pnpm parity:sync --all
|
|
pnpm parity:verify
|
|
# fix any agent-surface drift manually in the relevant agent/src/ files
|
|
git add . && git commit
|
|
```
|
|
|
|
### Adding a new instance
|
|
|
|
1. Create the new demo under `examples/integrations/<name>/` with a
|
|
Next.js frontend at the root and an `agent/` dir.
|
|
2. Add an entry to `manifest.json` under `instances`:
|
|
```json
|
|
"new-demo": {
|
|
"role": "instance",
|
|
"agent": { "language": "python", "runtime": "..." },
|
|
"allowedDivergence": ["agent/**", "src/app/api/copilotkit/**",
|
|
"Dockerfile", "docker/Dockerfile.agent",
|
|
"serve.py", "scripts/**"],
|
|
"packageJsonOverrides": { "scripts.dev:agent": "..." }
|
|
}
|
|
```
|
|
3. Run `pnpm parity:sync --target=new-demo`.
|
|
4. Hand-port the agent code (tools, state, prompt loading) — manifest
|
|
`tracked.agentSurface` tells you what tool names and state keys are
|
|
required.
|
|
5. Run `pnpm parity:verify --target=new-demo` until green.
|
|
|
|
### North-star's agent surface changed
|
|
|
|
Edit `manifest.json` → `tracked.agentSurface.toolNames` / `stateKeys`.
|
|
The verifier will then flag every instance that hasn't caught up.
|
|
|
|
### Canonical prompt changed
|
|
|
|
1. Edit `_parity/canonical/PROMPT.md`.
|
|
2. Update north-star's `agent/main.py` to use the new prompt string
|
|
(north-star is where humans read; canonical file is what the
|
|
verifier reads).
|
|
3. Port the new prompt into every instance's agent source (same
|
|
manual-merge rules as any agent change).
|
|
4. `pnpm parity:verify` — verifier greps first line of canonical against
|
|
each instance's agent source. Passes when all instances inline it.
|
|
|
|
## CI
|
|
|
|
`.github/workflows/integrations_parity.yml` runs `pnpm parity:check` on
|
|
every PR that touches `examples/integrations/**`. Failures link the
|
|
contributor back to this README.
|
|
|
|
## Design notes
|
|
|
|
- **Declarative, not prescriptive.** The manifest says _what_ is tracked;
|
|
the scripts just walk it. Adding a new tracked file = one line change,
|
|
not a code change.
|
|
- **Allowed-divergence is explicit, not implicit.** Everything is either
|
|
tracked, declared-divergent, or ignored — no silent "probably different"
|
|
state.
|
|
- **Agent surface is grep-level, on purpose.** A real AST check would be
|
|
three times the code and still miss semantic drift. The existing aimock
|
|
fixture integration tests (under `fixtures/default.json` per instance)
|
|
are the real correctness check; this is the "did someone rip out
|
|
`manage_todos`" safety net.
|
|
- **North-star is read-only to the scripts.** `sync.ts` refuses to write
|
|
into the north-star directory even if asked. Verifier never touches
|
|
anything.
|