Files
Alem Tuzlak 189eca6872 fix(showcase/ms-agent-python): modernize to gpt-5.2, native reasoning + multimodal, fix multi-turn A2UI
Removes obsolete workarounds and uses MAF's native primitives where the
framework supports them. Surfaces the genuine gaps as targeted shims.

Reasoning
  - delete the no-op `think` tool; configure `reasoning={effort,summary}`
    on the Agent default_options. AG-UI bridge already emits real
    REASONING_MESSAGE_* events from the Responses API.
  - bump `@ag-ui/client` 0.0.43 -> 0.0.52 so the frontend event schema
    includes REASONING_MESSAGE_* (0.0.43 was still on the deprecated
    THINKING_TEXT_MESSAGE_* union).
  - reasoning-block.tsx -> consume ReasoningMessage via the
    `messageView.reasoningMessage` slot; default-render demo strips
    its `useRenderTool(think)` and uses zero-config CopilotChat.

Multimodal
  - delete the frontend LegacyConverterShim that rewrote modern
    `{type:"image"|"document", source:...}` parts to legacy `binary`.
    MAF's AG-UI adapter handles the modern shape natively.
  - delete the pypdf extraction subclass; gpt-5.2 reads PDFs natively
    via OpenAI's `input_file`. Drop `pypdf` from requirements.
  - retain a tiny 30-line `_MultimodalAgent` + adapter monkey-patch
    that copies `metadata.filename` into Content.additional_properties
    so OpenAI's `input_file` requirement is satisfied (upstream gap
    in agent_framework_ag_ui._message_adapters._parse_multimodal_media_part).

A2UI dynamic
  - fix the flat ops shape bug in `build_a2ui_operations_from_tool_call`:
    middleware expects v0.9 nested `{createSurface:{surfaceId,catalogId}}`,
    not flat `{type:"create_surface",surfaceId,...}`. Flat shape silently
    fell back to surface group "default" and never rendered.
  - secondary structured-output call: use chat_client.client (underlying
    AsyncOpenAI) directly to bypass MAF's function-invocation auto-loop;
    inherits api_key + model from the parent. response_format=PydanticModel
    is unsuitable (strict-mode rejects open dicts); a one-shot raw-args
    primitive doesn't exist in agent_framework today.
  - inject the registered A2UI catalog schema (49KB of Zod types) from
    `input_data.context[]` into the secondary call's system prompt so the
    LLM emits correct prop names. _A2UIDynamicAgent captures the schema
    on each run().
  - switch the OUTER agent to OpenAIChatCompletionClient. Responses API
    + tool calls + multi-turn fails ("No tool output found for function
    call ...") because reasoning items can't be replayed; chat.completions
    replays tool history cleanly.
  - defensive _reorder_tool_messages shim: the CopilotKit frontend message
    store ships [user, tool, assistant(toolCalls), assistant, user] on
    turn 2 (tool BEFORE its parent assistant), which OpenAI rejects.
    Walk inbound messages and re-attach tool messages immediately after
    their matching assistant.toolCalls[].id. (Should be fixed upstream
    in @copilotkit/react-core/v2.)
  - tighten the render_a2ui JSON schema (minItems:1, explicit
    items.properties, strict:false) and the prompt header so gpt-5.x
    actually emits non-empty components with entry-level props.
  - same A2UI bypass replacement in agent.py + beautiful_chat.py (shared
    pattern, three call sites of the original `from openai import OpenAI`).

Model bump
  - default OPENAI_CHAT_MODEL_ID -> gpt-5.2 (Responses-capable reasoning
    model). Scoped clients: reasoning agent on gpt-5.2 + Responses;
    a2ui_dynamic on gpt-5.2 + chat.completions (avoids reasoning replay).

Upstream gaps surfaced (recommend follow-up PRs)
  - agent_framework_ag_ui: propagate part.metadata.filename to
    Content.additional_properties["filename"] in
    _parse_multimodal_media_part. (Removes the PDF subclass.)
  - agent_framework: a one-shot tool-args API
    (e.g. get_response(..., auto_invoke=False)) or a cleanly exported
    RawOpenAIChatClient. (Removes the A2UI raw-SDK bypass.)
  - @copilotkit/react-core/v2: stop reordering role=tool messages
    before their assistant parent in the frontend message store.
    (Removes the A2UI message-reorder shim.)
2026-05-12 15:06:11 +02:00

130 lines
4.5 KiB
Python

"""Dynamic A2UI tool: LLM-generated UI from conversation context.
This module provides the data preparation for a secondary LLM call that
generates v0.9 A2UI components. The actual LLM call is made by the
framework-specific wrapper (LangGraph, CrewAI, etc.) since each framework
has its own way of invoking LLMs.
"""
from __future__ import annotations
import logging
from typing import Any, Optional
_logger = logging.getLogger(__name__)
CUSTOM_CATALOG_ID = "copilotkit://app-dashboard-catalog"
# The render_a2ui tool schema that the secondary LLM is bound to.
RENDER_A2UI_TOOL_SCHEMA = {
"name": "render_a2ui",
"description": (
"Render a dynamic A2UI v0.9 surface.\n\n"
"Args:\n"
" surfaceId: Unique surface identifier.\n"
" catalogId: The catalog ID (use \"copilotkit://app-dashboard-catalog\").\n"
" components: A2UI v0.9 component array (flat format). "
"The root component must have id \"root\".\n"
" data: Optional initial data model for the surface."
),
"parameters": {
"type": "object",
"properties": {
"surfaceId": {"type": "string", "description": "Unique surface identifier."},
"catalogId": {"type": "string", "description": "The catalog ID."},
"components": {
"type": "array",
"items": {"type": "object"},
"description": "A2UI v0.9 component array (flat format).",
},
"data": {
"type": "object",
"description": "Optional initial data model for the surface.",
},
},
"required": ["surfaceId", "catalogId", "components"],
},
}
def generate_a2ui_impl(
messages: list[dict[str, Any]],
context_entries: Optional[list[dict[str, Any]]] = None,
) -> dict[str, Any]:
"""Prepare inputs for a secondary LLM call that generates A2UI components.
Returns a dict with:
- system_prompt: The system prompt for the secondary LLM (built from context)
- tool_schema: The render_a2ui tool schema to bind to the LLM
- tool_choice: The tool name to force
- messages: The conversation messages to pass through
- catalog_id: The default catalog ID
The framework wrapper should:
1. Make an LLM call with these inputs
2. Extract the tool call args (surfaceId, catalogId, components, data)
3. Build a2ui_operations from the args and return them
"""
context_text = ""
if context_entries:
context_text = "\n\n".join(
entry.get("value", "")
for entry in context_entries
if isinstance(entry, dict) and entry.get("value")
)
return {
"system_prompt": context_text,
"tool_schema": RENDER_A2UI_TOOL_SCHEMA,
"tool_choice": "render_a2ui",
"messages": messages,
"catalog_id": CUSTOM_CATALOG_ID,
}
def build_a2ui_operations_from_tool_call(args: dict[str, Any]) -> dict[str, Any]:
"""Build a2ui_operations dict from the secondary LLM's tool call args.
Call this after the framework wrapper extracts the tool call arguments.
"""
surface_id = args.get("surfaceId", "dynamic-surface")
catalog_id = args.get("catalogId", CUSTOM_CATALOG_ID)
components = args.get("components", [])
if not components:
_logger.warning("build_a2ui_operations_from_tool_call received empty components list")
data = args.get("data")
# A2UI v0.9 nested operation shape -- ``@ag-ui/a2ui-middleware`` reads
# ``op.createSurface.surfaceId`` / ``op.updateComponents.surfaceId`` to
# group activity events by surface. A flat
# ``{type: "create_surface", surfaceId, ...}`` shape silently parses
# (the middleware's ``getOperationSurfaceId`` returns ``undefined`` and
# falls back to "default") and the resulting activity event never
# matches a registered catalog surface, leaving a blank canvas.
ops = [
{
"version": "v0.9",
"createSurface": {"surfaceId": surface_id, "catalogId": catalog_id},
},
{
"version": "v0.9",
"updateComponents": {
"surfaceId": surface_id,
"components": components,
},
},
]
if data:
ops.append(
{
"version": "v0.9",
"updateDataModel": {
"surfaceId": surface_id,
"path": "/",
"value": data,
},
}
)
return {"a2ui_operations": ops}