## Summary
Automated refresh of the toolkit slugs the CLI knows without asking
the API, generated by
`ts/packages/cli/scripts/generate-toolkit-slugs.ts`.
Toolkits added since the last refresh currently cost users one
toolkit-list fetch (~2 s) the first time they run one of that
toolkit's tools. Merging this makes them free.
The generator refuses to write a list that is short, malformed, or
missing staple toolkits, so a bad fetch opens no PR at all.
## Summary
Automated sync of backend data into the docs site. Triggered by:
`schedule`.
## What changed
- **Toolkit catalog** (`docs/public/data/toolkits.json`,
`toolkits-list.json`) — refreshed list of available toolkits, auth
schemes, and tools from the backend API
- **OpenAPI specs** (`docs/public/openapi.json`,
`docs/public/openapi-v3.json`, `docs/public/openapi-webhooks.json`) —
latest v3.1 and v3.0 API specifications plus the webhook-events spec,
fetched from production
- **API reference pages** (`docs/content/reference/api-reference/`,
`docs/content/reference/v3/api-reference/`) — regenerated index pages
for both API versions
- **Meta tools reference** (`docs/public/data/meta-tools.json`,
`docs/content/toolkits/meta-tools/*.mdx`) — updated meta tool schemas
and reference docs
Co-authored-by: Sushmithamallesh <19796925+Sushmithamallesh@users.noreply.github.com>
One-line `ENGINE_REF` bump for the docs-agent-eval shim: the pin
predates the judge calibration (docs-agent-eval-ci PRs #4–#7 —
evidence-scoped scans, proxy-log ground truth, infra-vs-agent error
classification, corrected package taxonomy, renamed secret). Until this
merges, label/deployment-triggered evals run the old
false-positive-prone judge; dispatched runs already use current main.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Soumya Medapati <soumyamedapati@mac.local.meter>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This PR:
- supersedes https://github.com/ComposioHQ/composio/pull/4280,
https://github.com/ComposioHQ/composio/pull/4281, and
https://github.com/ComposioHQ/composio/pull/4282
- updates the Claude Code action, production dependencies, TypeScript
tooling, and experimental agent dependencies in one reviewed lockfile
refresh
- retains `@mastra/core@1.52.1` so Worker bundles do not pull in
Node-only `execa`
- aligns `ai@^7.0.79` with `@ai-sdk/mcp` and `@ai-sdk/openai` on
`@ai-sdk/provider-utils@5.0.30`
- adapts Eve approval-policy tests to its widened `ApprovalPolicy |
ApprovalConfiguration` contract
- adds the required patch changeset for the published
`@composio/experimental` TypeBox update
- verifies frozen install, package builds, workspace and example type
checks, all package tests, example validation, Oxlint, and Changesets
status
This PR:
- resolves internal `$ref`/`$defs` in tool input schemas before
translation in `langchain`, `llamaindex`, `claude-agent-sdk`, `vercel`,
`google`, and `openai-agents` (`onUnresolved: 'sentinel'`)
- previously `$ref`-typed properties degraded to `z.any()` — the Zod
converter has no `$ref` branch — or were emitted as dangling references
after the root rebuild (google, openai-agents fallback)
- keeps `openai-agents`' strict-structured-outputs path untouched:
OpenAI resolves `$defs`/`$ref` natively including recursion, pinned by a
guard test
- adds per-provider `$ref` regression suites for all six providers,
including dangling-`$defs` (`GMAIL_FETCH_EMAILS`) and recursive-schema
cases
- adds a cross-provider contract test that fails when a new provider
ships without a `$ref` classification, plus property tests for
`dereferenceJsonSchema` (Python counterparts land with the Python-side
fix)
- changes the vendor-visible schema shape for `$ref`-using tools;
changeset is `minor`
## Context
The same bug was fixed locally twice before (mastra, anthropic) without
surfacing the other six providers — nothing enumerated providers and
asked the `$ref` question. The new contract test does exactly that, so
provider #11 cannot ship unclassified. Python mirrors exist already; the
Python-side provider fix follows separately.
This PR:
- replaces the `console.log` in `wrapMcpServerResponse` with a redacted
`logger.debug` line (server names and count only)
- MCP URLs are user-scoped, bearer-equivalent endpoints; they previously
landed in stdout and aggregated production logs on every call
- adds a regression test asserting no stdout write and an unchanged
return shape
- treats MCP URLs already captured in existing logs as exposed;
regenerate those endpoints server-side
Regenerate the cross-provider $ref contract test, the
dereferenceJsonSchema property tests, and the openai-agents
$ref contract lost in a session handoff, ported from their
surviving Python counterparts. The openai-agents non-strict
ratchet flips to a plain it now that the fallback dereferences;
the superseded openai-agents-ref.test.ts is removed in favor of
the richer contract file.
jsonSchemaToZodSchema has no $ref branch, so a $ref node degrades to
z.any() for langchain, llamaindex, claude-agent-sdk, and vercel's
default path. google and openai-agents' non-strict fallback rebuild
the root from properties/required, discarding $defs while dangling
$ref pointers survive.
Dereference internal $ref/$defs before translation in all six, using
onUnresolved: 'sentinel' so a $ref into an undeclared $defs block
(e.g. GMAIL_FETCH_EMAILS) degrades to a permissive schema instead of
throwing. The mastra and anthropic providers already had this fix;
openai-agents' strict branch is untouched since OpenAI's structured
outputs support $defs/$ref natively, including recursion, and
google's rebuild still drops additionalProperties/title/root
oneOf-anyOf-allOf beyond the dangling-$ref class this fixes.
This PR:
- caps automatic S3 file downloads at 100 MiB in both SDKs — these were
the only fetch paths left without a size limit, and `s3url` is a
tool-execution response field, so the body behind it is no more trusted
than a user-supplied URL (the same input the SSRF guard already defends
against)
- TypeScript: `downloadFileFromS3` buffered the whole body with
`response.arrayBuffer()`; it now reads through the existing
`readResponseBodyWithLimit` guard, overridable per call via
`maxDownloadBytes`
- Python: `FileDownloadable.download` streamed straight to disk with no
byte accounting; it now uses the same `Content-Length` pre-check plus
authoritative streamed-byte counter as its sibling
`_fetch_file_from_url`, overridable via `max_size`
- fixes two latent defects in the Python write loop found along the way:
- an `OSError` from `fd.write` (disk full, permissions) escaped the
`ErrorDownloadingFile` contract the method documents, surfacing as a raw
`OSError(28)`
- every failure left a partial file on disk that no caller was told
about; cleanup now runs on all error paths via
`_discard_partial_download()`, which suppresses cleanup failures so an
`unlink` error cannot replace the error the caller needs to see
- adds 8 regression tests (2 TypeScript, 6 Python) covering oversized
`Content-Length`, oversized streaming with no/dishonest header,
partial-file cleanup, transport and write failures, and the within-limit
success path
- not a breaking change: the caps default to the existing 100 MiB used
by both sibling paths, and both new parameters are optional
## Context
The streamed byte counter — not the `Content-Length` check — is the
authoritative limit in both SDKs, because a malicious or misconfigured
server can omit or lie in that header. The header check is only an early
abort.
Deliberately out of scope: `RemoteFile.buffer()` / `RemoteFile.blob()`
in `ts/packages/core/src/models/RemoteFile.ts` are explicit
user-initiated whole-file reads. The same `readResponseBodyWithLimit`
treatment is the natural follow-up, but they are not on the automatic
path this PR closes.
Every other test in the package imports node builtins with the `node:`
prefix — this was the only bare `'fs'` specifier, and inconsistent with
`node:path` on the line above it.
Claude-Session: https://claude.ai/code/session_01K1hH9PMmd6KPKdkACX553z
Review follow-up on the download size cap.
`requests.exceptions.RequestException` subclasses `OSError`, so the two
handlers added for the write loop were byte-identical and the second already
subsumed the first. Collapse them into one `except OSError` and say why in a
comment, so the next reader does not re-add the redundant clause.
Route partial-file cleanup through `_discard_partial_download`, which
suppresses cleanup failures: an `OSError` from `unlink` would otherwise
replace the `ResponseTooLargeError` or transport error the caller needs.
Cover the two error paths that had no tests: a transport failure mid-stream
and a failing write both raise `ErrorDownloadingFile` and leave no partial
file behind. Without the handler the write failure escapes as a raw
`OSError(28)` — the defect these pin.
Claude-Session: https://claude.ai/code/session_01K1hH9PMmd6KPKdkACX553z
`FileDownloadable.download` streamed the response straight to disk with no
byte accounting, so an untrusted `s3url` could fill the disk. Add the same
`Content-Length` pre-check plus authoritative streamed-byte counter the
sibling `_fetch_file_from_url` already uses, capped at `_MAX_RESPONSE_SIZE`
and overridable per call via `max_size`.
Also close two gaps the write loop left open: an `OSError` from `fd.write`
(disk full, permissions) escaped the documented `ErrorDownloadingFile`
contract, and any failure left a partial file on disk that no caller was
told about. Every failure path now unlinks the partial file;
`ResponseTooLargeError` still propagates uncaught so callers see the limit.
Claude-Session: https://claude.ai/code/session_01K1hH9PMmd6KPKdkACX553z
`downloadFileFromS3` buffered the whole response with `arrayBuffer()`. The
`s3Url` it fetches is a tool-execution response field — the same untrusted
input the SSRF guard already defends against — so an oversized or endlessly
streaming body could exhaust the host process's heap.
Route the body through the existing `readResponseBodyWithLimit` guard, which
pre-checks `Content-Length` and counts streamed bytes (the header can be
absent or dishonest). The 100 MiB default matches the upload-from-URL sibling
in the same module; `maxDownloadBytes` overrides it per call.
Claude-Session: https://claude.ai/code/session_01K1hH9PMmd6KPKdkACX553z
This PR:
- unblocks https://github.com/ComposioHQ/composio/pull/4177 by
documenting TypeScript `@composio/core` `0.18.0`
- bumps Python `composio` and all 12 provider distributions from
`0.20.0` to `0.21.0`
- keeps `python/composio/__version__.py` and the pinned `uv.lock`
aligned with package metadata
- adds the combined customer-facing changelog for strict tool schemas,
safer file transfers, and runtime reliability updates
- records the Node.js 22.22.3 minimum for the TypeScript release
## Release sequence
1. Merge this PR into `next`.
2. Merge #4177 after its release-workflow check turns green; Changesets
publishes the TypeScript packages to npm.
3. Tag the resulting `next` commit as `py@0.21.0` to publish the Python
core and provider packages to PyPI.
## Verification
- `pnpm test:release-workflow`
- `mise exec -- uv lock --check`
- `make chk`
- `make build`
- `uv tool run twine check dist/*` (26 artifacts)
- `bun run test` in `docs/` (527 tests)
## What
`ComposioError` (`errors/ComposioError.ts`) assigns `name` as a class
field (`public name = 'ComposioError'`). Under the package tsconfig
(es2022 -> `useDefineForClassFields`), each subclass must reassign
`this.name` in its constructor or it inherits the base value. ~30
sibling subclasses do this; three omitted it:
- `ComposioToolVersionRequiredError` (`errors/ToolErrors.ts`)
- `JsonSchemaToZodError` (`errors/ValidationErrors.ts`)
- `JsonSchemaRefResolutionError` (`errors/ValidationErrors.ts`)
So `new JsonSchemaToZodError().name === 'ComposioError'`, and the same
for the other two. All three are thrown on real paths (`Tools.ts`,
`jsonSchema.ts`), and `error.name` is forwarded to error telemetry
(`telemetry/Telemetry.ts`), so these distinct error types silently
mis-group under the base name; any consumer branching on `err.name ===
'<ClassName>'` never matches.
## Fix
Add the missing `this.name = '<ClassName>'` at the end of each of the
three constructors, matching the established sibling pattern.
Runtime-only; no type or public-API change. (The two `PusherErrors`
subclasses set `name` via a class field, which already resolves
correctly, so they are intentionally left untouched.)
## Tests
Adds `test/errors/errorNames.test.ts` asserting each of the three
reports its own class name and is `instanceof ComposioError`. Verified
fails-before / passes-after; full `@composio/core` suite green (1095
tests), plus `tsc`, oxlint, and prettier clean.
## What
`TelemetryService.sendMetric` and `sendErrorLog` issue `await
fetch(...)` with no timeout. If the telemetry endpoint stalls, the await
never settles. Both are awaited on the SDK's telemetry path
(`Telemetry.ts` batch-processor callback, `sendMetric`, and
`sendErrorTelemetry`), so a stalled telemetry endpoint can leave an SDK
call pending indefinitely — which the existing `catch` comment says must
never happen ("telemetry failures should never affect SDK calls").
## Fix
Wrap both requests in a private `postWithTimeout` helper that bounds
each best-effort request with an `AbortController` + `setTimeout` (3s)
and clears the timer in `.finally()`. This mirrors the already-merged
bound on the background npm version check in `utils/version.ts` (#4027),
including the hand-rolled-timer-over-`AbortSignal.timeout` rationale: an
uncleared timer pins the workerd request context open for the full
timeout on every successful send. On timeout the abort rejects `fetch`,
which the existing `catch` swallows exactly as it already swallowed
network errors, so the best-effort / never-throws contract is unchanged.
## Tests
Adds a `TelemetryService network bounding` suite (mirrors
`version.test.ts`): asserts each method passes an `AbortSignal` and
clears its timer on success, and that a never-responding endpoint
resolves to `undefined` without throwing. Verified fails-before /
passes-after; full `@composio/core` suite green (1092 tests), plus
`tsc`, oxlint, and prettier all clean.
This PR:
- builds on top of https://github.com/ComposioHQ/composio/pull/4209 by
@AseemPrasad, keeping its recursive `toStrictJsonSchema()` and Python
parity while changing the mechanism so strict mode stops deleting
parameters
- keeps every optional parameter under `strict: true`: properties become
required and optional ones accept `null` (the emulation OpenAI
documents), instead of dropping 42% of parameters across the 930-tool
corpus; `type` arrays stay as they are, so nullable objects stay
nullable
- sends tools whose schema strict mode cannot express (objects with
arbitrary keys, `allOf`, `prefixItems`, dangling `$ref`s, non-object
roots) with `strict: false` and a warning naming the tool and path,
instead of narrowing them to empty closed objects
- adds `omitNullToolArguments()`: the strict providers drop a `null`
argument only where the tool's own schema rejects it, so nullable fields
keep an explicit `null`
- keeps local `$ref`/`$defs` (recursion included) under strict mode
instead of inlining them
- brings the Python `OpenAIResponsesProvider(strict=True)` to parity:
emits `strict`, calls the base initializer, mirrors the rewrite and null
omission
- makes Mastra and openai-agents use the same strict semantics
(`OpenAIAgentsProvider({ strict: true })` previously had no effect)
- pins the behavior with a shared `strict-cases.json` corpus
(byte-identical TypeScript/Python copies) plus edge-case regression
tests enumerated independently with a second model
## Context
OpenAI's structured-outputs contract accepts `"type": ["string",
"null"]` and rejects `type` next to `anyOf`. Validated against OpenAI's
own `toStrictJsonSchema` converter over the 930 real tool schemas in
`ts/packages/cli/test/__mocks__/tools.json`: the previous approach was
accepted for 930/930 tools but only after removing 1,682 properties;
this one emits strict schemas for 816 tools (all accepted, no property
lost, idempotent) and downgrades the 114 tools that use free-form or
map-style parameters.
https://claude.ai/code/session_01TDrxCHn2hg51HmxVstSUgs
## Summary
Auto-generated TypeScript SDK reference docs from
`ts/packages/core/src/`.
Regenerates pages at `docs/content/reference/sdk-reference/typescript/`
to reflect changes in the core package's public API (new methods,
updated signatures, changed types).
## Summary
Automated sync of backend data into the docs site. Triggered by:
`schedule`.
## What changed
- **Toolkit catalog** (`docs/public/data/toolkits.json`,
`toolkits-list.json`) — refreshed list of available toolkits, auth
schemes, and tools from the backend API
- **OpenAPI specs** (`docs/public/openapi.json`,
`docs/public/openapi-v3.json`, `docs/public/openapi-webhooks.json`) —
latest v3.1 and v3.0 API specifications plus the webhook-events spec,
fetched from production
- **API reference pages** (`docs/content/reference/api-reference/`,
`docs/content/reference/v3/api-reference/`) — regenerated index pages
for both API versions
- **Meta tools reference** (`docs/public/data/meta-tools.json`,
`docs/content/toolkits/meta-tools/*.mdx`) — updated meta tool schemas
and reference docs
## Summary
- stop installing the independently distributed Autogen adapter into the
shared CrewAI/LangChain/LangGraph unit-test environment
- run the Autogen import guard and signature regressions in their own
matrix environment
- align the local `tst` nox session with the compatible shared provider
set and add `tst_autogen` for isolated Autogen coverage
## Root cause
`autogen-core==0.7.5` requires `protobuf~=5.29.3`, while CrewAI's
current telemetry dependency chain requires
`googleapis-common-protos>=1.75.1`, whose generated modules require
`protobuf>=6.33.5`.
The packages are independently distributed adapters and are not tested
together, but the workflow installed both into one virtual environment.
Because the installs were sequential, installing Autogen last downgraded
`protobuf` to `5.29.6` and left the already-installed Google modules
unusable:
```text
google.protobuf.runtime_version.VersionError: Detected incompatible Protobuf Gencode/Runtime versions when loading google/rpc/error_details.proto: gencode 6.33.5 runtime 5.29.6.
```
## Regression coverage
The shared suite intentionally skips Autogen because loading it
alongside CrewAI creates the incompatible protobuf environment. CI now
runs these existing regressions in the isolated Autogen environment
instead:
- `test_autogen_signature_honors_skip_defaults`
- `test_autogen_signature_preserves_default`
Developers can reproduce that boundary with `nox -s tst_autogen`.
## Verification
- reproduced the downgrade after the Autogen provider installation
- shared provider environment imports `google.rpc.error_details_pb2`
with `protobuf==6.33.6`
- isolated Autogen environment imports `composio_autogen` with
`protobuf==5.29.6`
- isolated Autogen regressions: 2 passed
- Python unit suite: 1,355 passed, 35 skipped
- Ruff and mypy: passed
- agent-skill and skill-routing validators: passed
- Prettier and `git diff --check`: passed
No Changeset is required: this only changes CI and test-environment
setup.