Commit Graph

1257 Commits

Author SHA1 Message Date
Ran Shemtov fa13d52502 Merge branch 'main' into codex/crewai-full-d6 2026-08-14 09:37:42 +02:00
Mark f97f0768ba test(showcase): isolate CrewAI resume bridge contracts
Exercise both bridge bindings without leaking monkeypatches, and verify rejected bridge versions cannot mutate either binding.
2026-08-13 16:46:26 -07:00
Mark 2116257e1e test(showcase): harden CrewAI cancellation regressions
Use bounded dispatch and cancellation waits in both CrewAI integrations, and verify any fallback worker finishes during cleanup.
2026-08-13 16:46:14 -07:00
Mark 35aa2a34a0 fix(showcase): close CrewAI cancellation edge cases
Use AsyncOpenAI so cancellation reaches the in-flight GenerateA2UI request while retaining the thread fallback for synchronous backend tools.

Preserve cancelled versus resolved-null interrupts across pinned ag-ui-crewai 0.3.0 by encoding only resolved null as JSON null and failing loudly on version drift.

Reuse the canonical shared render_a2ui schema so the secondary request remains aligned with the shared tool contract.
2026-08-13 16:46:05 -07:00
Mark 8a6d14b29a fix(showcase): port pydantic-ai integration to v2 and restore live system prompts (#6379)
Ports `showcase/integrations/pydantic-ai` — the last pydantic-ai surface
still on v1 — to Pydantic AI v2. Refs #6364.

Three commits plus a bot formatting fix, best reviewed separately.

## 1. `chore(showcase): port pydantic-ai integration to Pydantic AI v2`

- **`requirements.txt`** → `pydantic-ai-slim[ag-ui,openai]==2.22.0`,
`ag-ui-protocol==0.1.19`. Drops the `opentelemetry-api<1.44` ceiling
from #6374; v2 resolves cleanly against otel 1.44.0, so the workaround
is no longer needed. `starlette<1.0.0` is unchanged and satisfies v2's
`>=0.46.2`.
- **9 `StateDeps` imports** move from `pydantic_ai.ag_ui` (removed in
v2) to `pydantic_ai.ui`.
- **`agent_server.py`** — `Agent.to_ag_ui()` was removed in 2.0.0, so a
`mount_agent()` helper builds the equivalent Starlette sub-app and
mounts it. The shape is deliberately identical to what v1's `AGUIApp`
produced — a Starlette app whose only route is `POST /`, named
`run_agent` — so **all 19 mount paths behave exactly as before, trailing
slashes included, and no TypeScript route file changes**.

`deps` is constructed **per request**. v1's `run_ag_ui` did `deps =
replace(deps, state=state)`, handing each run its own object; v2's
adapter does `deps.state = state`, mutating what it is given. A single
shared instance under v2 therefore lets concurrent runs overwrite each
other's state mid-run.

## 2. `fix(showcase): apply the multimodal provider gate to v2 native
content`

v2's `AGUIAdapter.load_messages` converts AG-UI attachments to native
content types *before* the model boundary; v1 delivered the raw AG-UI
part dicts. `_NATIVE_CONTENT` listed `BinaryContent` as a flatten
fixpoint, so under v2 inline attachments were waved straight through and
the entire provider gate was skipped:

- inline PDFs were no longer text-extracted, so raw bytes went to OpenAI
- unsupported image subtypes (HEIC/SVG/TIFF) were no longer degraded and
reached the provider as images, which fails the turn
- missing-mime magic-byte sniffing never ran
- `AudioUrl`/`VideoUrl` were neither fixpoints nor classifiable, so they
hit the fail-loud raise

`BinaryContent` is no longer a fixpoint. `_classify_native_content` maps
native content onto the same `(kind, scheme, mime, value)` tuple the
AG-UI classifier already produces, so **every existing gate applies
unchanged** — no gate logic was rewritten. `audio/*` and `video/*` are
named explicitly because `_kind_for` routes them to `"other"`, and a
missing mime defaults to `"image"` so the sniffer runs.

Net behaviour matches v1: a supported inline image still flattens to an
`ImageUrl` data URI, which is why most of the suite went green without
touching assertions.

Five assertions did change. They checked that state-backing content was
still AG-UI `InputContent`, which encoded v1's bridging. They now assert
the flatten's output (`ImageUrl`) never appears in state — the leak they
were written to guard. The adjacent identity and snapshot checks that
prove non-mutation are untouched.

## 3. `fix(showcase): gate url-source content and correct the v1-parity
claim`

Adversarial review of the first two commits found the gate was only half
fixed. `_NATIVE_CONTENT` still short-circuited `ImageUrl` and
`DocumentUrl`, which v2 builds from unvetted client input, so url-source
attachments bypassed the gate where v1 routed them through it:

- an `image/heic` or `image/svg+xml` url reached the provider as
`input_image`, which the Responses API rejects — failing the turn
- an `audio/mpeg` document url reached it as `input_file`
- a blank-mime inline PDF went to the image sniffer instead of text
extraction, because `load_messages` collapses `ImageInputContent` and
`DocumentInputContent` to the same bare `BinaryContent` and erases the
modality v1 defaulted on

Native content is now gated **before** the fixpoint check rather than
instead of it. `_classify_native_content` returns a tuple only when the
gate must act; `None` means provider-safe and falls through to the
fixpoint, preserving object identity. `ImageUrl` is gated rather than
rerouted so a provider-safe one keeps its identity and any explicit
`_media_type`.

It also corrected a false claim. The `mount_agent` docstring said
routing *and* behaviour were unchanged. Routing is; model input is not.
v2 defaults `manage_system_prompt='server'`, so each agent's
`system_prompt=` now reaches the model. On v1 it never did —
`_agent_graph` emitted system parts only `if not messages` and the AG-UI
bridge always supplied history — so **18 of 19 agents had silently dead
system prompts on main**. A/B on both versions with the same agent and
request: v1 sends 0 system-prompt parts, v2 sends 1. The new behaviour
is correct and kept; the docstring now says so.

## Verification

Against pydantic-ai 2.22.0, in a venv built from this branch's
`requirements.txt`:

- **52/52 Python tests pass**, up from 42/52.
`test_multimodal_content_mapping.py`'s `importorskip` pointed at the
removed `pydantic_ai.ag_ui`, which would have skipped all 43 of its
tests **green** under v2; it now targets `pydantic_ai.ui.ag_ui` and uses
the public `AGUIAdapter.load_messages` in place of the v1 private
helper.
- **16/19 mounts** return `200 text/event-stream` with `RUN_STARTED …
RUN_FINISHED` and no `RUN_ERROR`, driven through the real app with
`TestClient` using trailing-slash URLs as the TS routes do. The other
three (`/a2ui_dynamic`, `/beautiful_chat`, `/`) reach tool execution and
then fail on a raw `OpenAI()` client constructed inside a tool, which
the harness cannot intercept and aimock handles in CI.
- **Per-request deps isolation** confirmed on
`/shared_state_read_write`: state sent by one request does not appear in
the next.

`build-check (pydantic-ai)` is green on this branch, and because
`requirements.txt` changed, the cached pip layer was invalidated — so
that was a **genuine fresh resolve of pydantic-ai 2.22.0 inside the real
Dockerfile**, not a cached pass. It also confirms dropping the
`opentelemetry-api<1.44` ceiling is safe.

### D6 harness probes — run, with a baseline

The behavioural gate is the shared harness D6 probes. No CI job runs
them for showcase paths, so they were run locally on both this branch
and `main`:

| | main (v1) | this branch (v2) |
|---|---|---|
| passed | **33** / 36 | **34** / 36 |
| `reasoning-display` | ✗ `no reasoning-role message rendered within
5000ms` | ✅ **passes** |
| `gen-ui-agent` | ✗ `waitForTurnComplete … runStartCount=2,
done-signal-missing` | ✗ identical error |
| `shared-state-read` | ✗ `Strict mode: 1 candidate fixture(s) skipped
by sequence/turn state` | ✗ identical error |

```bash
cd showcase
AIMOCK_URL_LOCAL=http://localhost:4010 bin/showcase test pydantic-ai --d6 --direct --rebuild --cycle --verbose
```

**The port takes D6 from 33/36 to 34/36.** The two remaining failures
are pre-existing on `main` with byte-identical error strings — this
branch neither causes nor fixes them, and both are tracked in #6381
rather than blocking here.

`gen-ui-agent` is root-caused and is not fixture drift: that demo was
never ported to pydantic-ai. `src/agents/gen_ui_agent.py` exists in
llamaindex with a real `set_steps` tool but has no counterpart here, the
route points at `/gen_ui_tool_based/` (the chart-viz agent), and
`set_steps` is declared nowhere in the package. The fixture fabricates
`set_steps` calls the backend cannot honour, so pydantic-ai rejects the
unknown tool and exhausts its single retry. Confirmed live against real
OpenAI: the cell returns plain text, which is correct for the code as
written.

`reasoning-display` going green is the notable behavioural gain, and it
retires a documented v1 limitation. `PARITY_NOTES.md:91-97` justifies
omitting the reasoning-message branch of `use-rendered-messages.tsx` on
the grounds that "PydanticAI's AG-UI adapter does not emit reasoning
content today" — true on v1, false on v2. (That block is stale on two
further counts: it cites `@ag-ui/core@0.0.43` where `package.json` pins
0.0.57, and claims `ReasoningMessage` is not exported where it is
imported at `reasoning-block.tsx:4`.) Correcting it is tracked on #6364.

To be precise about what that proves: **v2 forwards reasoning content
where v1 dropped it.** The probe supplies the reasoning channel via its
fixture, so what is verified is the forwarding path — adapter → AG-UI
stream → frontend renderer — end to end. Whether a given model actually
emits a reasoning summary live is a separate matter and outside this
port's control: it requires a native reasoning model
(`reasoning_agent.py` defaults to `gpt-5`, overridable via
`REASONING_MODEL`) and, for summary text, a verified OpenAI
organisation. A live run here returned prose with no reasoning block,
consistent with the org-verification gate rather than anything in the
port.

Also verified: the image builds from scratch on v2. Because
`requirements.txt` changed, the cached pip layer was invalidated, so
`build-check (pydantic-ai)` in CI was a genuine fresh resolve of
pydantic-ai 2.22.0 inside the real Dockerfile — which also confirms
dropping the `opentelemetry-api<1.44` ceiling is safe.

### CI gate coverage, for the record

No CI job exercises this package's runtime behaviour on a PR, on this
branch or on `main`:

- `test / e2e / dojo` runs from the upstream `ag-ui` checkout (`ref:
main`) against upstream example agents, and filters on `packages/**` /
`sdk-python/**`
- `test_showcase-frontend-matrix.yml` is dispatch-only and builds the
integration from `base/` — a frozen-backend React baseline
- `showcase_validate.yml` asserts `tests/e2e/` exists with a minimum
spec count; it does not run it
- the package's own `tests/e2e/` (37 files) is invoked by nothing — per
`AGENTS.md` rule 1 the measuring test is the shared harness probe, so
that layer is legacy

## Remaining for #6364

Two acceptance criteria are outstanding, which is why this says Refs
rather than Closes:

- the harness D6 value-test (`bin/showcase test pydantic-ai --d6
--rebuild`), which no CI gate runs for showcase paths
- `PARITY_NOTES.md` has 6 version-dependent blocks, 4 of which were
already inaccurate against the tree before this PR; left alone
deliberately to keep this diff scoped

## Possible follow-up

`multimodal_agent.py` still reaches into three private APIs
(`pydantic_ai._run_context`, `pydantic_ai.models.wrapper`,
`pydantic_ai.models.{ModelRequestParameters,StreamedResponse}`) and
subclasses `WrapperModel`, overriding
`request`/`count_tokens`/`request_stream`. v2 adds a supported
alternative: `AbstractCapability.before_model_request`, which receives a
`ModelRequestContext` carrying `messages` and `streaming`. Migrating
would delete those private imports and ~85 lines. Deliberately not in
this PR — it fixes nothing and would obscure the review.
2026-08-13 09:59:40 -07:00
Ran Shem Tov 48a01b6203 fix(showcase): harden CrewAI probe parity 2026-08-13 00:10:36 +02:00
Ran Shem Tov 79f56d9f8e fix(showcase): finalize CrewAI D6 on official bridge 2026-08-11 22:45:04 +03:00
Ran Shem Tov 6862508eb2 Merge remote-tracking branch 'origin/main' into codex/crewai-full-d6
# Conflicts:
#	showcase/harness/Dockerfile
#	showcase/scripts/fail-baseline.json
2026-08-07 17:53:55 +03:00
Ran Shem Tov 61eed4a4ad fix(showcase): harden CrewAI D6 parity on a3 2026-08-07 17:50:13 +03:00
Alem Tuzlak 6f640f7eb1 fix(showcase/ms-agent-dotnet): surface shared-state-read-write chat replies (#6233)
## Summary

`shared-state-read-write` pills showed **no chat responses** on staging.

### Cause

#6227 wired deterministic replies for the suggestion pills, but those
updates were emitted as:

```csharp
new AgentRunResponseUpdate { Contents = [new TextContent(...)] }
```

without `Role = ChatRole.Assistant`. AG-UI's .NET adapter only turns
assistant-role text into `TEXT_MESSAGE_*` events, so the frontend
dropped every pill reply. Notes snapshots could still land; chat looked
dead.

### Fix

- Set `Role = ChatRole.Assistant` on deterministic text updates
- Prefer `message.Text` when resolving the latest user message
- Broaden pill matching for greet / weekend / remember-something copy

## Test plan

- [x] `dotnet build` ms-agent-dotnet agent
- [ ] Staging after deploy: Greet / Remember something / Plan a weekend
all show assistant text; Remember something updates the notes panel
2026-08-07 16:31:02 +02:00
Alem Tuzlak d5d2e73a53 fix(showcase/ms-agent-dotnet): ground declarative-gen-ui charts in sales data (#6232)
## Summary

`declarative-gen-ui` on staging painted surfaces but charts showed **No
data available** and tables were empty.

### Cause

With `injectA2UITool: false`, the secondary design LLM does **not**
receive frontend App Context (`useSalesAnalystContext` /
sales-context.ts). It only got a thin design prompt, so it omitted or
emptied `PieChart`/`BarChart` `data` arrays and `DataTable` rows.

### Fix

- Embed the Vantage Threads Q2 dataset + composition rules into
`DeclarativeGenUiDesignSystemPrompt`
- Add concrete non-empty PieChart / BarChart / DataTable examples
- Coerce string chart values to numbers
- Tighten outer agent: one short sentence, no prose dashboards

## Test plan

- [x] GenerateA2ui unit tests 12/12
- [ ] Staging after deploy: all four declarative-gen-ui pills show
populated charts/tables from the Q2 dataset
2026-08-07 16:29:19 +02:00
Mark 0c10d8c882 docs(pydantic-ai): remove duplicate quickstart, fix dead links and commands
Mechanical repairs found while auditing the pydantic-ai docs. Each was
verified against the tree; nothing here is a content rewrite.

- Delete `quickstart/pydantic-ai.mdx` + its `meta.json`. `seo-redirects.ts`
  already routes `/pydantic-ai/quickstart/pydantic-ai` ->
  `/pydantic-ai/quickstart` (rule F6), and adk got the same treatment (F7).
  pydantic-ai was the only framework still carrying a `quickstart/`
  subdirectory alongside the canonical `quickstart.mdx`.
- `human-in-the-loop/agent.mdx`: link to the canonical quickstart directly
  instead of the redirected legacy path, and point the starter link at
  `examples/integrations/pydantic-ai` — `examples/coagents-starter-pydantic-ai`
  does not exist.
- `docs-links.json`: `subagents.shell_docs_path` was `/multi-agent/subagents`,
  which has no page. The real page is `/multi-agent-flows`, which the
  entry's own `og_docs_url` already pointed at.
- `headless-simple/chat.tsx`: the console tag said `langgraph-python` inside
  the pydantic-ai package. This sits in an `@region` block, so it is pulled
  into docs as a snippet. 11 other integrations carry the same copy-paste;
  they are left for the fleet sweep.
- `examples/showcases/pydantic-ai-todos/README.md`: `uv run src/main.py` ->
  `uv run main.py` (there is no `src/main.py` in that tree), and the stated
  Python floor now matches `agent/pyproject.toml` (`>=3.13`).
- `examples/canvas/pydantic-ai/README.md`: Python 3.8+ was unrunnable —
  `agent/agent.py` uses PEP 604 unions. Aligned to the sibling tree that
  pins the same `pydantic-ai-slim==2.22.0`.
2026-08-06 21:47:23 +00:00
Mark 3cf128ec8f docs(showcase): correct pydantic-ai v2 API refs and gen-ui-agent comments
PARITY_NOTES.md and qa/beautiful-chat.md described `agent.to_ag_ui()`,
which v2 removes. Replaced with the AG-UI adapter / `mount_agent()`
wording this branch introduces.

Also flags the PARITY_NOTES "Skipped demos" section as stale rather than
silently leaving it: mcp-apps, hitl-in-chat and hitl-in-chat-booking all
ship, and the reasoning/interrupt reasons no longer match manifest.yaml
(which is the authority). Full rewrite tracked in OSS-777.

The gen-ui-agent comments asserted a `src/agents/gen_ui_agent.py` and a
`set_steps` tool that exist nowhere in this package. The cell has no route
override, so it proxies to the root sales agent and its D6 probe is red on
main (GH #6381). The comments now describe that, instead of an
intended-but-unbuilt contract.

Adds the missing `shared-state-read` entry to manifest.yaml `demos:` — it
was declared under `features:` with no route or highlight. Mirrors
langgraph-python's entry, which likewise omits an agent file because the
cell runs on the neutral default agent.
2026-08-06 21:42:05 +00:00
Mark dc3681b106 Merge branch 'main' into chore/showcase-pydantic-ai-v2 2026-08-06 10:11:29 -07:00
Ran Shem Tov e7cc29bfc0 docs(showcase): consolidate conversational flows under CrewAI 2026-08-06 16:40:01 +03:00
Ran Shem Tov 0b6129141c docs(showcase): document CrewAI CF version floor 2026-08-06 15:45:27 +03:00
Ran Shem Tov 5136097aa0 feat(showcase): add CrewAI conversational flows 2026-08-06 15:33:10 +03:00
Ran Shemtov 9b768a0b98 feat(showcase): finalize MAF Python - D6 green on agent-framework 1.0 latest (#5985)
## Finalize MAF Python: D6 green on official agent-framework 1.0 latest

Brings the `ms-agent-python` showcase integration to a clean,
reproducible D6 state on the officially published latest
`agent-framework` packages, with feature parity to `langgraph-python` on
everything buildable today.

### Dependencies (exact pins, official latest)

- `agent-framework-ag-ui==1.0.1`
- `agent-framework-openai==1.12.0`
- `agent-framework-core==1.13.0`

No beta/rc floors, no ranges. Removed two unused `langchain-*` deps. All
framework deps are exact pins; `validate-pins` ratchet baseline moves
down 31 to 27. The only remaining ms-agent-python pin FAIL is the
shared-frontend `openai ^5.9.0`, identical across every integration
(pre-existing baseline).

### D6 result: all green on the published mock

Verified with `showcase test ms-agent-python --d6 --direct --rebuild`
against the actual published `ghcr.io/copilotkit/aimock:latest`
(**v1.38.0**), freshly pulled: 37 distinct cells executed, 37
conversations completed, zero failures, aggregate `d6:ms-agent-python
green (104.2s)`.

`tool-rendering-reasoning-chain` (previously the only red on the
published mock) is now green: it needed `reasoning.encrypted_content`
echoed back on the second Responses request (upstream
microsoft/agent-framework#7233), which the published mock did not
synthesize until
[aimock#342](https://github.com/CopilotKit/aimock/pull/342), shipped in
aimock **v1.38.0**. Fixed upstream, not worked around.

`multimodal` is un-quarantined and now matches langgraph. It had been
wrongly marked unsupported based on a local-only failure: the
`sample.png`/`sample.pdf` demo assets are Git LFS pointers, and without
git-lfs on PATH the attachment send fails before the run starts
(`runStartCount=0`). langgraph-python multimodal fails locally for the
identical reason yet declares the feature supported. Verified the MAF
agent works (D6 cell green with the real assets, 2 turns, assertions
passed); both production deploys serve the real 10KB PNG.
`not_supported_features` now equals langgraph exactly:
`[gen-ui-interrupt, interrupt-headless]` (both a shared
`@copilotkit/react-core/v2` resume-path bug, quarantined in langgraph
too).

### Cells fixed on this branch

- `tool-rendering-custom-catchall` (18-entry fixture +
MESSAGES_SNAPSHOT-drop subclass so narration renders last)
- `shared-state-streaming` (seed `/document` after RUN_STARTED +
`chunkSize` fixtures so replay emits per-token deltas)
- `tool-rendering-reasoning-chain` (un-quarantined; green on aimock
v1.38.0)
- `frontend-tools-async` (removed a stray broad fixture that
shadowed/looped)
- `open-gen-ui` + `open-gen-ui-advanced` (removed six stray fixtures
colliding in the shared gen-ui fixture file)
- `multimodal` (un-quarantined; parity with langgraph)

### Deferred to upstream (not worked around)

- **a2ui-recovery**: langgraph ships a bespoke A2UI validate-and-retry
recovery demo. MAF Python's A2UI is going native via
[microsoft/agent-framework#7423](https://github.com/microsoft/agent-framework/pull/7423),
which delivers progressive streaming, error recovery, and the sub-agent
design built into `agent-framework-ag-ui`, and even includes the same
two bridge fixes hand-rolled here (unanswered-tool-call stripping + A2UI
MESSAGES_SNAPSHOT suppression). Building a bespoke recovery demo now
would be throwaway. When #7423 merges and releases, the showcase A2UI
migrates to the native path and the recovery demo lands with it.

### Validators

- `generate-registry`: OK
- `validate-pins`: 27 fails, hash matches ratcheted baseline
- `validate-parity`: PASS
- `validate-fixture-tool-surface`: clean

### Notes

- `useCoAgent` is deprecated; all demos use `useAgent` from
`@copilotkit/react-core/v2`.
- Kept in draft pending review. No blocking external gates: aimock#342
shipped in v1.38.0.
2026-08-06 12:13:09 +02:00
Ran Shemtov a7006cd1ed Merge branch 'main' into claude/elated-snyder-01a8ce 2026-08-06 08:09:08 +02:00
Ran Shem Tov ab94c1315e fix(showcase): let the Mastra MCP Apps agent self-correct a rejected diagram
Switching models only moved the failure rate around, it never removed it, so
stop relying on the model getting hand-escaped JSON right on the first try.

`create_view` takes `elements` as a stringified JSON array. When the model
appends a stray `}` past the closing `]`, the MCP server rejects the call and
names the exact fault ("Invalid JSON in elements: Unexpected non-whitespace
character after JSON at position N"). That error already comes back as a tool
result, and the agent had no step cap, so a retry was mechanically possible
all along. What blocked it was our own prompt: "Call create_view ONCE" and
"do NOT iterate, do NOT make multiple calls. Ship on the first shot."

The prompt now tells the model to read the error and try again, capped at 2
corrections (3 calls total), with stopWhen: stepCountIs(6) bounding the loop
if it never converges. This mirrors the validate-then-retry recovery pattern
already used for A2UI on the other integrations.

Validated against the real Excalidraw MCP server, using the agent's prompt
extracted verbatim from this file and the real tool schema:

  normal runs                       12/12 succeeded, all on the first call
  attempt 1 force-corrupted with
  the real-world stray `}`          10/10 recovered on the second call

Also verified in the running app (local dev server, real key): valid JSON,
isError false, diagram rendered.

Not yet verified in-app: the recovery path itself. No natural failure occurred
during the in-app runs, so the retry is proven at the API level rather than
through the Mastra agent loop.
2026-08-05 22:55:49 +03:00
Ran Shem Tov 88d9719faf docs(showcase): connect CrewAI full parity docs 2026-08-05 22:38:46 +03:00
Ran Shem Tov ccf979eca8 fix(showcase): stabilize remaining CrewAI D6 cells 2026-08-05 22:38:23 +03:00
Ran Shemtov 8f24b0373b Merge branch 'main' into claude/framework-d6-integration-validate-7be45e 2026-08-05 21:08:41 +02:00
Ran Shemtov f58cc22750 Merge branch 'main' into claude/competent-chatelet-7305ee 2026-08-05 20:53:04 +02:00
Ran Shem Tov c60233acb2 fix(showcase): close CrewAI D6 tool lifecycles 2026-08-05 19:15:33 +03:00
Ran Shem Tov 20f2ff8f4c fix(showcase): move Mastra MCP Apps agent to gpt-5.4
Owner preference for the 5.x line. Recorded honestly: this reduces the
empty-diagram failure but does not remove it.

Measured against the real Excalidraw MCP server (same system prompt, real
tool schema, via the Responses API the AI SDK actually uses):

  gpt-4o-mini   create_view OK 3, isError 5
  gpt-5.4       create_view OK 7, isError 3
  gpt-4.1       create_view OK 8, isError 0
  gpt-5.5       create_view OK 10, isError 0

JSON validity of the `elements` argument:

  gpt-4o-mini   7 invalid of 12
  gpt-5.4       7 invalid of 28, plus two runs whose tool call came back
                garbled with unrelated spam text
  gpt-5.4 + a hardened prompt   2 invalid of 16 (prompting does not fix it)
  gpt-4.1       0 invalid of 12
  gpt-5.5       0 invalid of 16

So roughly 30% of diagrams still render as an empty iframe on gpt-5.4. Closing
that gap needs a follow-up, most likely validating or repairing the `elements`
string before the MCP call rather than relying on the model to hand-escape
nested JSON correctly.
2026-08-05 19:02:11 +03:00
Ran Shem Tov 08c38950da fix(showcase): route CrewAI D6 native flows 2026-08-05 18:55:40 +03:00
Ran Shemtov fed66e046f Merge branch 'main' into ran/pni-121-mastra-tool-rendering-results-delivered-out-of-sequence 2026-08-05 17:41:33 +02:00
Ran Shemtov 73fe341c3d Merge branch 'main' into claude/competent-chatelet-7305ee 2026-08-05 17:41:18 +02:00
Ran Shem Tov b1aebc706d fix(mastra-showcase): match gold model — beautifulChatAgent on gpt-5.4
Gold `beautiful_chat.py` uses `ChatOpenAI(model="gpt-5.4")`; align the mastra
beautifulChatAgent to the same model (was gpt-4o). Verified live: dashboard turn
stays [query_data, generate_a2ui] (no standalone chart over-call), full surface
renders (3 metrics + pie + bar).
2026-08-05 18:38:28 +03:00
Ran Shem Tov adb23bfdc6 fix(showcase): use local CrewAI recovery context 2026-08-05 18:36:33 +03:00
Ran Shem Tov 2aa6f35edf fix(showcase/mastra): align tool-rendering agent to gold's model
PNI-121's reported symptom was "the weather appears after 'Find flights from
SFO to JFK'". Two separate things produced that screenshot; the blank flight
rows were the previous commit. This is the other one.

On gpt-4o this agent opened the flights turn by restating the PREVIOUS turn's
weather ("The weather in San Francisco is currently 20C with heavy rain.")
before narrating the flights, which reads exactly like a tool result arriving
a turn late. Nothing was actually late: only one weather card exists and it
stays in turn 1, and the flights turn's card holds the flights turn's data.
The system prompt is already byte-identical to gold tool_rendering_agent.py,
so the model was the remaining divergence - gold runs gpt-5.4.

Verified live (real LLM, ticket's exact click order - "Weather in SF" then
"Find flights"): the stale weather sentence is gone and the narration now
matches gold's shape ("SFO -> JFK options: United UA231 08:15-16:45 $348; ..."
against gold's "Flights SFO -> JFK: United UA231 08:15-16:45 $348; ...").
Weather, stock and d20 pills re-checked on the same rig.

No effect on CI: aimock fixtures never match on model, and the d20 pill's
5-roll chain is fixture-scripted (7/14/3/19/20), so the e2e sequence is
unchanged. Live, gold rolls the d20 once too - so this moves the demo toward
gold rather than away from it.
2026-08-05 18:32:59 +03:00
Ran Shem Tov 801913c801 feat(showcase): enroll CrewAI in all 41 D6 cells 2026-08-05 18:32:31 +03:00
Ran Shemtov a04c708e5d Merge branch 'main' into claude/framework-d6-integration-validate-7be45e 2026-08-05 17:31:21 +02:00
Ran Shem Tov 753b1445eb fix(mastra-showcase): stop beautiful-chat dashboard from over-calling standalone charts + align flights fixture
Issue: asking for the Sales Dashboard (esp. after a prior turn) made the model
call the standalone pieChart + barChart frontend tools AND generate_a2ui, so
loose charts painted next to the dashboard.

- beautifulChatAgent: mirror gold `beautiful_chat.py` `parallel_tool_calls=False`
  (defaultOptions.providerOptions.openai.parallelToolCalls) + sharpen the
  steering so a dashboard / "using A2UI" request calls generate_a2ui ONLY (it
  draws the charts inside the surface), while a single-chart request still uses
  the standalone pieChart/barChart tool. Verified live: dashboard turn now calls
  [query_data, generate_a2ui] only, single + after-flights.

- aimock beautiful-chat flights fixture: model the fixed-schema `search_flights`
  path (returns the A2UI FlightCard envelope directly) instead of the old
  generate_a2ui -> render_a2ui chain, so the fixture matches the live behavior
  and the e2e spec's stated intent (United $349 / Delta $289).
2026-08-05 18:23:35 +03:00
Ran Shem Tov 8f74bcb738 feat(showcase): add CrewAI state and multimodal flows 2026-08-05 18:22:36 +03:00
Ran Shem Tov 02c0daf28a fix(showcase): settle CrewAI interrupt resumes 2026-08-05 18:16:51 +03:00
Ran Shem Tov 1265509f37 feat(showcase): add native CrewAI Flow interrupts 2026-08-05 17:51:43 +03:00
Ran Shem Tov 36f05f17b9 fix(showcase): render Mastra MCP Apps diagrams reliably on live
The MCP Apps cell intermittently rendered an empty iframe (an empty box or a
thin band) on a live endpoint while passing under aimock.

Root cause is the model, not the renderer. Excalidraw's `create_view` declares
`elements` as `type: "string"` holding a JSON array, so the model must emit
double-encoded, hand-escaped JSON. gpt-4o-mini frequently appends a stray `}`
just past the closing `]`. The MCP server then rejects the call with
"Invalid JSON in elements: Unexpected non-whitespace character after JSON",
returns isError, and there is no diagram to draw, so the iframe paints empty.
aimock never catches this because it replays a recorded, valid payload.

Measured against the real Excalidraw MCP server using this agent's exact
system prompt and the real tool schema:

  gpt-4o-mini   create_view OK 3, isError 5
  gpt-4.1       create_view OK 8, isError 0

And on JSON validity alone (n=12 unless noted):

  gpt-4o-mini                 8 invalid
  gpt-4o-mini + hardened prompt 7 invalid  (prompting does not fix it)
  gpt-4.1-mini                3 invalid, and it bloats output
  gpt-4.1                     1 invalid of 24, output stays compact

gpt-4.1 is already used by five other agents in this file, so this keeps the
integration consistent. The cell's aimock fixture does not key on the model, so
d6 replay is unaffected.
2026-08-05 17:43:45 +03:00
Ran Shem Tov 33fca30cad fix(showcase/mastra): return gold-shaped search_flights result so flight rows render
On a live endpoint the tool-rendering flight card rendered every row blank
("United ? -> ? --") while the model's narration below it carried the real
times and prices. The result was delivered in full; the card just never
matched it.

search_flights emitted Mastra-flavored keys (flightNumber / departureTime /
arrivalTime / price) but FlightListCard - in both this integration and gold
langgraph-python, byte-identical - reads { airline, flight, depart, arrive,
price_usd }, which is exactly what gold's tool_rendering_agent.py returns.
Only `airline` overlapped, so the rest fell back to placeholders.

Return the gold result shape directly. The legacy caller-supplied `flights`
passthrough is untouched, and the only consumers of this tool are the three
tool-rendering-style agents, all of which drive gold-shaped cards.

Verified on a live real-LLM endpoint (no aimock): reproduced the blank rows
before the change, then confirmed all three rows render
"United UA231 08:15 -> 16:45 $348" (plus Delta and JetBlue) after it, across
tool-rendering, tool-rendering-custom-catchall and
tool-rendering-reasoning-chain, and across a two-turn weather + flights
conversation.

The existing e2e only asserted origin/destination and a row count, so blank
rows passed. It now asserts the result's depart/arrive/price and rejects the
"? -> ?" placeholder.

Fixes PNI-121
2026-08-05 17:35:36 +03:00
Ran Shem Tov 7e85d964d5 feat(showcase): use native CrewAI reasoning and inputs 2026-08-05 17:35:35 +03:00
github-actions[bot] 8df39dd9bb style: auto-fix formatting 2026-08-05 14:33:00 +00:00
Ran Shem Tov d9bf3efcf1 fix(mastra-showcase): render beautiful-chat flights as fixed-schema A2UI (langgraph-python parity)
The "Search Flights (A2UI Fixed Schema)" pill narrated flight results as plain
text instead of rendering FlightCards. Root cause: mastra's beautiful-chat
reused the shared `searchFlightsTool` (returns plain `{ flights }`, which the
tool-rendering cells render via their own frontend FlightListCard), so nothing
produced an A2UI surface and the model just described the data.

langgraph-python's beautiful_chat.py wires a DEDICATED fixed-schema
`search_flights` whose tool RESULT is a complete `a2ui_operations` envelope (a
flat Row of literal FlightCards on `app-dashboard-catalog`, surface
`flight-search-results`). Mirror it:

- Add `searchFlightsA2uiTool` returning that envelope (buildFlightCardComponents
  mirrors `_build_flight_components` — inline values, no structural-template
  children).
- Add a dedicated `beautifulChatAgent` (query_data, todos, generate_a2ui, the
  fixed search_flights, + the flight/dashboard steering prompt) and point the
  beautiful-chat route at it, so the fixed-schema flights and steering don't
  leak into the shared weatherAgent / tool-rendering cells.

Verified live (real LLM, dedicated agent): the flights pill paints two
FlightCard surfaces (catalogId app-dashboard-catalog, surface
flight-search-results), and the Sales Dashboard dynamic surface still renders in
full (metrics + pie + bar) via the grounded generate_a2ui.
2026-08-05 17:30:55 +03:00
Ran Shem Tov 33aa7fed82 chore(showcase): validate CrewAI alpha stack 2026-08-05 17:26:38 +03:00
Ran Shem Tov aa15946c2a fix(mastra-showcase): ground beautiful-chat A2UI dynamic render from request context
The dynamic `generate_a2ui` tool grounded its inner `render_a2ui` subagent from
the tool's `contextEntries` arg, which the outer model always sends empty
(captured live: `{"messages":[…],"contextEntries":[]}`). On a live LLM the inner
render then ran with an EMPTY system prompt: ungrounded, it emitted
invalid/misnamed components (or none), so the surface never resolved against the
catalog. Result: the Beautiful Chat A2UI dynamic surface (Sales Dashboard,
flights) rendered no UI, with a render error that varied run to run. aimock hid
it: the recorded fixture returns a valid envelope regardless of the empty
context.

Read the catalog schema + A2UI generation guidelines the `@ag-ui/mastra` bridge
already forwards onto Mastra's request context
(`requestContext.get("ag-ui").context`) and ground the render there instead of
trusting the model-supplied arg. Mirrors `readAgUiContext` in `@ag-ui/mastra`'s
`getA2UITools`. Preserves per-demo catalogId (the grounded model emits it, and
`buildA2uiOperationsFromToolCall` uses `args.catalogId`), so the same shared
weatherAgent serves both beautiful-chat and declarative-gen-ui. Falls back to
the arg when no request context is present.
2026-08-05 14:20:44 +03:00
Ran Shem Tov 528ebfa351 fix(showcase): un-quarantine MAF Python multimodal (parity with langgraph)
multimodal was wrongly quarantined based on a local-only failure: the
sample.png/sample.pdf demo assets are Git LFS pointers, and without
git-lfs on PATH the attachment send fails before the run starts
(runStartCount=0). langgraph-python multimodal fails locally for the
exact same reason, yet declares the feature supported.

Verified the MAF agent actually works: with the real assets in place
the D6 cell is green (2 turns, assertions passed). Both production
deploys serve the real 10KB PNG (LFS pulled), so multimodal is green
in prod for MAF just like langgraph. not_supported_features now matches
langgraph exactly: [gen-ui-interrupt, interrupt-headless].
2026-08-05 14:12:16 +03:00
Ran Shem Tov 47c56c8f9c fix(showcase): MAF Python manifest + exact pins for clean validate
- Remove multimodal from features (already in not_supported_features);
  the duplicate made generate-registry reject the manifest, blocking D6
  from starting.
- Pin agent-framework-{ag-ui,openai,core} to exact published latest
  (1.0.1 / 1.12.0 / 1.13.0) so validate-pins classifies them as exact
  framework deps; drop unused langchain-openai/langchain-core.
- Ratchet validate-pins baseline down 31 -> 27 (the 4 non-exact framework
  fails are now exact); ms-agent-python's only remaining FAIL is the
  shared-frontend openai ^5.9.0, identical to every other integration.
2026-08-05 14:12:16 +03:00
Ran Shem Tov a71768d3c4 chore(showcase): MAF Python deps floor to explicit agent-framework 1.0+ latest
Replace main's ugly beta floors (ag-ui 1.0.0b251117 / openai 1.0.0rc6) with
clean 1.0+ ranges. Resolves to latest (ag-ui 1.0.1, core 1.13.0, openai 1.12.0);
D6 36/36 green on that.
2026-08-05 14:12:16 +03:00
Ran Shem Tov 02f079180d fix(showcase): MAF Python frontend-tools-async + gen-ui-open green; quarantine multimodal
- frontend-tools-async: drop the stray broad 'project planning' fixture entry that
  substring-shadowed 'Find my notes about project planning' -> run-loop. Green.
- gen-ui-open + gen-ui-open-advanced: remove 6 gen-ui-open-owned strays from
  gen-ui-tool-based.json ('3D axis visualization', 'Inline expression evaluator',
  'render an open gen-ui element', 'continue the advanced gen-ui flow') that
  collided (same context, different responses) with gen-ui-open.json. Green;
  gen-ui-tool-based unaffected.
- multimodal: quarantine. On main + latest deps the browser gets 200 from the
  runtime but never starts a run (runStartCount=0) — a shared multimodal
  frontend/runtime issue (identical frontend to langgraph; agent works via direct
  POST; my #5985 fix doesn't change it). Genuine gap, needs shared-layer work.

Full D6 now 36/36 green on official latest (ag-ui 1.0.1, core 1.13.0, openai 1.12.0).
2026-08-05 14:12:16 +03:00
Ran Shem Tov b56bb182d2 feat(showcase): MAF Python shared-state + reasoning-chain green on main (latest deps)
- shared-state-streaming: replace main's stale 'counter' fixture (drifted from
  langgraph canonical) with the write_document document demo (6-entry poem/email/
  quantum + chunkSize) + SharedStateStreamingFrameworkAgent seed subclass; frontend
  already matches langgraph. Un-quarantine + add to features. Green (3/3 turns).
- tool-rendering-reasoning-chain: un-quarantine + add to features. Green on core
  1.13.0 / openai 1.12.0 (latest) via aimock encrypted_content (CopilotKit/aimock#342,
  which fixes the Responses reasoning multi-tool regression) + the store:False agent.
  CI-green depends on aimock#342 releasing.
- Drop reasoning-default-render / agentic-chat-reasoning from not_supported (no D6
  probe featureType — not real cells).
2026-08-05 14:12:16 +03:00