## Summary
`composio --version`: 288ms to 199ms. Peak RSS: 97.8MB to 77.3MB.
Executable: 85.9MB to 79.7MB. Every command benefits.
A compiled Bun binary parses its whole embedded bundle before running
any JS. #4468 stopped the TypeScript compiler and the tokenizer rank
table from being evaluated at startup, but they were still parsed every
time. The compiler was 44% of the executable's JavaScript, the o200k
table another 28%. Both now ship as files next to the executable and
load on demand.
Fourth PR in the stack. Stacked on #4468; review #4463, #4464 and #4468
first. #4475 builds on this one.
Bun 1.4.1+4661e494f, linux-x64, best of 15, telemetry disabled, both
binaries built in the same session:
| | before (#4468) | after |
|---|---|---|
| `composio --version` | 288ms | 199ms |
| `composio tools execute --help` | 287ms | 202ms |
| peak RSS | 97.8MB | 77.3MB |
| executable | 85.9MB | 79.7MB |
| executable JS, minified | 8.3MB | 2.1MB |
Across the whole stack, from `next`: `--version` 612ms to 184ms, peak
RSS 167MB to 78MB, executable 95.8MB to 79.7MB.
`composio execute` end to end, against the live backend with a logged-in
CLI, best of 7 for the small response and best of 5 for the large one.
Tool: `HACKERNEWS_GET_ITEM_WITH_ID` (no connected account needed) and
`HACKERNEWS_GET_LATEST_POSTS`. "Tail" is the time from the
`execute.tool_call.end` perf event to process exit.
| | `next` | #4468 | this PR |
|---|---|---|---|
| 1.6KB response, wall | 2431ms | 1857ms | 1761ms |
| 1.6KB response, tail | 294ms | 12ms | 11ms |
| 35KB response, wall | 2665ms | 2210ms | 2165ms |
| 35KB response, tail | 322ms | 329ms | 353ms |
The stack removes ~670ms from a small execute: ~430ms of startup and
~280ms of tokenizer construction that no longer happens. The large
response keeps its ~330ms tail because past 10KB the tokenizer is still
built; this PR adds ~20ms there for the on-demand parse of the encoder
file. The remaining ~1.7s is network the stack does not touch: DNS and
TLS to the backend, the preflight round trips before `tool_call.start`,
and the session create plus execute pair. Wall times move by ±150ms
between runs because of that; the tail column is the stable one.
## Changes
1. `generation-runtime.mjs` carries `src/generation/*`, the `composio
run` source rewrites, `typescript`, `@composio/ts-builders` and
`openapi-typescript`. `generate ts`, `generate py` and `run` load it
with the new `loadInstalledCompanionModule`. From a source checkout the
loader resolves the `.ts` file next to `run-companion-modules.ts`, so
tests and `bun run src/bin.ts` need no build step. The specifier is
computed at runtime on purpose; a literal `import('./x')` gets folded
back into the executable. Before importing, a packaged install runs the
self-repair download only if that companion's own files (its wrapper and
what the wrapper imports) are missing, so a different missing file
cannot block it. The repair has to come first, because Bun keeps a
failed or already-loaded import in its module registry. A file that
fails to import, or lacks one of the exports its caller names, is a
typed `RunCompanionRepairError` asking to reinstall, not a crash. Both
companions are also tsdown entries, so the `dist/` build resolves them.
2. `execute-output-encoder-runtime.mjs` carries `js-tiktoken/lite` plus
the rank table. `execute` loads it only past the 10KB byte gate from
#4463, and never for executes started by `composio run`. If it cannot be
loaded, even after the self-repair download, `execute` estimates the
token count from the byte length (about four bytes per token) instead of
failing a tool call that already succeeded. The estimate can undercount,
so such a response is always stored as a file rather than printed
inline.
3. Both join `RUN_COMPANION_MODULE_BASENAMES`, the mechanism `composio
run` already uses for its helpers, so build, release packaging, install
and upgrade verification, and the self-repair download pick them up
unchanged. The three hand-maintained uninstall lists and the upgrade E2E
fixture gain the two file names.
4. A companion bundles its own copy of `effect`, and a fiber cannot run
primitives from another copy. So nothing Effect-shaped crosses the
boundary. The generation companion exposes plain promises and returns
failures as values. `src/generation/errors.ts` rebuilds them as the
CLI's own error classes with fields and stack intact.
5. `src/constants.ts` imported `constants` from `@composio/core`'s root
entry for two strings and two URLs, which evaluated the whole SDK at
startup (~25ms, mostly zod schemas). The values are inlined and a test
pins them to core's. `tool-file-uploads.ts` imports its three core
helpers on the upload path instead of at module scope.
6. Build guard. After building the companions, the build bundles
`src/bin.ts` once more with the release build's `DEBUG_OVERRIDE_*` env
inlining, defines, `NODE_ENV=production` and syntax minification
(whitespace is kept, so the per-module path comments it reads survive),
and fails if the executable's graph reaches `typescript`, `js-tiktoken`,
`src/generation/*` or a companion entry. Checked that it fires on a
stray static import. `@composio/core`'s root entry is not on the list:
it is still bundled behind the file-upload path's dynamic import (see
Additional context), so the guard cannot exclude it.
`test/src/commands/startup-imports.test.ts` forbids the same modules
when the command tree loads from source.
What changes for users:
- A damaged install (companion file missing) now affects `generate` the
way it already affected `run`: self-repair from the release archive,
then an error. A large `execute` also attempts the repair, and if that
fails it stores the response with a byte-based token estimate rather
than failing. Responses under 10KB never touch the encoder. `--version`
and everything else are unaffected.
- `composio upgrade` from a binary older than this PR copies only the
companion files that binary knows about. The first `generate`, `run` or
large `execute` on the new version then restores the two new files
through the self-repair download.
- `execute` responses over 10KB pay ~20ms more after
`execute.tool_call.end` (351 to 374ms), the on-demand parse of the 2.2MB
encoder file. Under 10KB, unchanged.
- Errors from generation are rebuilt instances. Same class, tag, fields,
message and stack; different object identity.
Generated output is byte-identical to #4468 for `generate ts`, `generate
ts --transpiled` and `generate py`. The 11-invocation help/error diff
from #4468 is identical.
Found on the way: `assertBundledRuntimeFiles` blanked string literals to
same-length runs of spaces, and the import patterns' `^\s*` then
backtracked quadratically over the compiler's embedded lib strings. The
build hung for over ten minutes. String bodies are dropped now. The
check has also never matched a specifier, since the specifiers it looks
for are the strings it removes. Left as is, because a corrected version
flags false positives in `run-subagent-output-mcp`.
## Type of change
- [ ] Bug fix
- [ ] New feature
- [x] Refactor/Chore
- [ ] Documentation
- [ ] Breaking change
## How Has This Been Tested?
Bun 1.4.1+4661e494f, Node 24.20.0, pnpm 11.8.0, linux-x64.
1. `cd ts/packages/cli && pnpm run typecheck && pnpm run
validate:boundaries && pnpm run validate:skills`
2. `pnpm exec vitest run`: 129 files, 1335 passed, 1 skipped. New tests
cover the mirrored constants, error rehydration and outcome lifting, and
the loader resolving both companions from source.
3. `pnpm build:binary`, then against `dist/composio`: `generate ts`,
`generate ts --transpiled` and `generate py` diffed against #4468's
binary, `run` with a trailing expression, `execute` with 1.6KB and 35KB
responses, and the damaged-install cases with files deleted from
`dist/`.
4. Docker E2E on this branch: `upgrade` 2 pass, `run` 8 pass, `version`
9 pass, `install` 7 pass on bash and 5 pass on zsh. The install runs
used a fixture built the way CI builds it (`build:binary:cross`,
`build:binary:package`, `build:binary:checksums`), which also confirms
the release zip carries both new files.
5. `bun run test/release-workflow.test.ts` at the repo root, for the
synced uninstall lists.
6. Follow-up commit (encoder fallback, typed load failure, graph-check
and tsdown fixes): `pnpm run typecheck`, `validate:boundaries` and
oxlint pass. The execute, companion-loader, constants,
generation-runtime, `run` and `generate` suites pass (177 passed, 1
skipped), including new tests for the estimate when the encoder cannot
load and for the typed load failure. The fallback test fails without the
fix. `pnpm build` emits both companions, and `pnpm build:binary` passes
the graph check.
7. Review follow-ups (companion repair scoped to the requested module
and run before the import, required-export check, stored output when the
token count is an estimate, graph check using the release build inputs,
startup-imports list): `pnpm run typecheck`, prettier and oxlint pass.
Full `pnpm exec vitest run` on the CLI package: 1340 passed, 1 skipped,
1 failure in `analytics.dispatch.test.ts`, which this PR does not touch
and which fails 1 run in 3 on its own. After the last loader change, the
companion-loader, execute, `run`, `generate`, generation-runtime,
startup-imports and upgrade suites pass (198 passed, 1 skipped). The new
fallback test for a response whose estimate is under the threshold fails
without its fix. `bun run ./scripts/build-companion-modules.ts` passes
the updated graph check.
## Screenshots (if applicable)
Not applicable.
## Checklist
- [x] I have read the Code of Conduct and this PR adheres to it
- [x] I ran linters/tests locally and they passed
- [x] I updated documentation as needed
- [x] I added tests or explain why not applicable
- [ ] I added a changeset if this change affects published packages
`@composio/cli` is private, so no changeset. Docs: the uninstall lists
and the code generation section of `ts/packages/cli/AGENTS.md`.
## Additional context
`openai` and `pusher-js` (~0.5MB minified) are still in the executable.
Only `@composio/core`'s root entry reaches them, and the two upload
guards have no lighter subpath export. A
`@composio/core/utils/file-upload-guard` entry would remove them; that
is a core package change.
Companion files carry no version stamp. The loader checks that a
companion has every export its caller uses, so a file from another
release missing one fails with a reinstall error. A file from another
version with the same exports still loads as is; checking `APP_VERSION`
after import would catch that, and it is not done here.
The remaining ~95ms of module evaluation is a long tail of eager Schema
and command definitions across `src/commands`, `src/services`, `effect`
and `src/models`, not one dependency.
Bun keeps a failed or already-loaded import in its module registry, so
importing a companion again after a repair can still see the missing
dependency or the old module. Repair now runs before the single import, and
only when the requested companion's wrapper or its import graph is missing,
so unrelated missing companions still cannot block it.
Claude-Session: https://claude.ai/code/session_01MTb47vexN35pLGsSZwBmrJ
- Import an in-process companion before checking the rest of the set, so a
missing unrelated companion cannot fail generate or run through an offline
repair. Repair runs only when the requested module fails to load.
- Check each companion's required exports, so a file left by another release
fails with a typed reinstall error instead of calling a missing export.
- When the tokenizer cannot load, store any response past the byte pre-filter.
The four-bytes-per-token estimate can undercount, so it no longer keeps a
response inline.
- Bundle the executable graph check with the release build's env inlining,
defines, NODE_ENV and syntax minification.
Claude-Session: https://claude.ai/code/session_01MTb47vexN35pLGsSZwBmrJ
Resolve the generate and run loaders in favor of the generation companion, drop the entry modules it supersedes, and let the startup-imports test allow src/generation/errors.ts as the binary build guard does.
Claude-Session: https://claude.ai/code/session_01MTb47vexN35pLGsSZwBmrJ
## Summary
`composio --version`: 622ms to 408ms. Eager module evaluation: 364ms to
130ms.
`commands/index.ts` builds the root command tree from every `.cmd.ts`,
so evaluating one command evaluated all of them. Two of them reached the
TypeScript compiler and the code generation pipeline at module scope.
`composio execute` paid ~165ms for a compiler it never called.
Stacked on #4464. Review #4463 and #4464 first.
Bun 1.4.1+4661e494f, linux-x64, best of 7, analytics disabled, same
script before and after:
| | before | after |
|---|---|---|
| `composio --version` | 622ms | 408ms |
| module evaluation | 363.8ms | 130.0ms |
| `commands/run.cmd` | 155.8ms | 8.0ms |
| `commands/generate` | 63.5ms | 2.5ms |
## Changes
`Command.withHandler` runs lazily, so moving an import inside a handler
body defers it. Specs, flags, descriptions and subcommand wiring still
resolve eagerly, so parsing, help and "did you mean" suggestions cannot
change.
1. `run.cmd.ts` was the only consumer of `import ts from 'typescript'`,
through three source rewrites `composio run` applies to a user script.
They move to `run-source-transforms.ts`, which the handler imports
dynamically. Tests import from the new path.
2. `ts.generate.cmd.ts` and `py.generate.cmd.ts` pulled
`src/generation/*` at module scope. Both resolve it inside the handler
now, right before first use.
These use `Effect.promise`, not `Effect.tryPromise`. A rejected import
of a module bundled into this binary is a broken build, not a
recoverable failure.
## Type of change
- [ ] Bug fix
- [ ] New feature
- [x] Refactor/Chore
- [ ] Documentation
- [ ] Breaking change
## How Has This Been Tested?
Bun 1.4.1+4661e494f, Node 24.17.0, pnpm 11.8.0, linux-x64.
1. Built the binary before and after and diffed stdout, stderr and exit
code across 11 invocations: `--help` at root and for generate, generate
ts, generate py, run, tools and execute, plus `version`, `--version`, an
unknown command and an unknown flag. Identical. The error paths are
there on purpose; they exercise the parser and the suggestion code,
where a shifted tree would show first.
2. `pnpm run typecheck && pnpm run validate:boundaries && pnpm run
validate:skills`
3. `pnpm test`: 1326 passed, 1 skipped, 1 failed. The failure is
`test/src/cli-main.test.ts`, which spawns the CLI from source against a
15s timeout and takes ~24s in this container. It fails the same way on
the parent commit (25.6s and 25.2s there, 24.5s and 24.3s here).
Reproduce: `cd ts/packages/cli && pnpm build:binary && time
./dist/composio --version`.
After rebasing onto the updated #4463 and #4464: `pnpm run typecheck`
passes, and the `run`, `generate ts`, `generate py` and `execute` suites
pass (120 passed, 1 skipped). The code in this PR is unchanged.
## Screenshots (if applicable)
Not applicable.
## Checklist
- [x] I have read the Code of Conduct and this PR adheres to it
- [x] I ran linters/tests locally and they passed
- [ ] I updated documentation as needed
- [ ] I added tests or explain why not applicable
- [ ] I added a changeset if this change affects published packages
No docs describe module loading order. No new tests; the existing suite
covers the moved functions, and the 11-invocation diff covers what this
could break. A test asserting the module is not loaded eagerly would be
good to have; #4469 adds a build-time check instead. `@composio/cli` is
private, so no changeset.
## Additional context
~130ms of eager evaluation remains. `services/agents` is 98ms of it:
Effect `Schema` definitions built at module scope. It cannot be deferred
as-is because `effects/handle-agent-auth-error.ts` narrows with `error
instanceof AgentAuthError` and six handlers depend on it. That is a
separate change.
The ~235ms pre-main bundle parse is unaffected. It scales with bundle
size, and a dynamic import keeps the module in the bundle. A binary that
bundles everything but runs only `console.log` still costs ~235ms. #4469
moves the code out of the bundle.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01EzaE7oGVgziJ5nRvBhcci2
The hung-native-host test forked setup, yielded once, then advanced the
TestClock by two minutes. Setup does real file I/O before it reaches the
host command, so on a slow CI worker the timeout was not yet armed when
the clock moved; the sleep then waited forever and the test failed on
vitest's own 15s limit. It flaked on next-based branches today, not
only this one. The hanging runner now opens a latch when setup reaches
it, and the test advances the clock only after that.
Claude-Session: https://claude.ai/code/session_01MTb47vexN35pLGsSZwBmrJ
A static import of typescript or src/generation anywhere on the command
tree silently restores ~165ms of module evaluation to every invocation,
and nothing failed. The test loads the command tree in a fresh Bun
process and asserts, via the module registry, that neither the compiler
nor the generation pipeline nor the run source transforms were
evaluated. Verified it fails on a stray static import.
Claude-Session: https://claude.ai/code/session_01MTb47vexN35pLGsSZwBmrJ
The generate handlers assembled their pipeline from three or four dynamic
imports and repackaged the results by hand. Each pipeline now has an
index module that re-exports what the handler needs, so the deferred
load is a single import and the boundary the handler crosses is named.
Claude-Session: https://claude.ai/code/session_01MTb47vexN35pLGsSZwBmrJ
A large execute measured its output after the tool call had already succeeded,
so a missing encoder companion whose repair download failed turned a successful
call into a failed command. Fall back to a byte-based token estimate instead.
Load companion modules with Effect.tryPromise so an unloadable file is a typed
RunCompanionRepairError rather than a defect, and add the two companions as
tsdown entries so the dist build can resolve them.
The executable graph check listed @composio/core's root entry with a pattern
that could never match Bun's relative module paths. The root entry is still
bundled behind the file-upload dynamic import, so drop that entry and correct
the constants comment.
`composio --version` goes from 288ms to 199ms, peak RSS from 97.8MB to
77.3MB, and the executable from 85.9MB to 79.7MB. Every command benefits.
A compiled Bun binary parses its whole embedded bundle before the first
line of JavaScript runs, and #4468 had already made sure the TypeScript
compiler and the tokenizer rank table were never *evaluated* unless
`generate`, `run`, or a large `execute` response needed them. They were
still *parsed* on every start: the compiler alone was 44% of the
executable's JavaScript and the o200k rank table another 28%, so
`--version` spent ~75ms reading code it could never call.
Both now ship as companion modules next to the executable, through the
mechanism `composio run` already uses for its own runtime helpers:
- `generation-runtime.mjs` carries `src/generation/*`, the `composio run`
source rewrites, `typescript`, `@composio/ts-builders` and
`openapi-typescript`. `generate ts`, `generate py` and `run` load it
with `loadInstalledCompanionModule`; from a source checkout the loader
resolves the `.ts` next to `run-companion-modules.ts` instead, so tests
and `bun run src/bin.ts` need no build step.
- `execute-output-encoder-runtime.mjs` carries `js-tiktoken/lite` and the
rank table. `execute` loads it only once a response exceeds the 10KB
byte pre-filter.
A companion bundles its own copy of `effect`, and a fiber cannot run
primitives built by another copy of the runtime, so nothing Effect-shaped
crosses the boundary: the generation companion exposes plain functions
and promises, runs its pipelines on its own runtime, and returns failures
as values that `src/generation/errors.ts` rebuilds as the CLI's own error
classes, stack included. Generated output is byte-identical to #4468 for
`generate ts`, `generate ts --transpiled` and `generate py`.
Both modules join `RUN_COMPANION_MODULE_BASENAMES`, so the build, release
packaging, install verification, `upgrade` and the self-repair download
pick them up unchanged. The three hand-maintained uninstall lists and the
upgrade E2E fixture gain the two file names.
Two smaller startup costs go with it:
- `src/constants.ts` imported `constants` from `@composio/core`'s root
entry for two strings and two URLs, which evaluated the whole SDK at
startup (~25ms of module-scope work, mostly zod schemas). The four
values are spelled out and pinned to core's by a test.
- `tool-file-uploads.ts` imported three core helpers at module scope that
only a file upload reaches; they are imported on that path now.
The binary build gains a guard: after bundling the companions it bundles
`src/bin.ts` once more unminified and fails if the executable's graph
reaches `typescript`, `js-tiktoken`, core's root entry, `src/generation/*`
or a companion entry. Without it a stray static import would put the
compiler back into the executable with nothing to notice.
Building also surfaced that `assertBundledRuntimeFiles` blanked string
literals to same-length runs of spaces, which made the import patterns'
`^\s*` backtrack quadratically across the compiler's multi-megabyte
embedded lib strings and stalled the build for over ten minutes. String
bodies are dropped now. (The check itself has never matched a specifier,
since the specifiers it looks for are the string literals it removes;
that is left as it was.)
Measured on the pinned toolchain, Bun 1.4.1+4661e494f, linux-x64, best
of 15, telemetry disabled, both binaries built in the same session:
composio --version 288ms -> 199ms
tools execute --help 287ms -> 202ms
peak RSS 97.8MB -> 77.3MB
executable 85.9MB -> 79.7MB
executable JavaScript 8.3MB -> 2.1MB (minified)
The `execute` tail after `execute.tool_call.end` is unchanged for
responses under 10KB (~10ms) and ~20ms slower above it (351 -> 374ms),
which is the on-demand parse of the 2.2MB encoder companion.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wx9gEjuiHux2weiHjdNcDs
`composio --version` drops from 622ms to 408ms, and eager module evaluation
from 363.8ms to 130.0ms, measured on the pinned toolchain (Bun
1.4.1+4661e494f, linux-x64, best of 7, analytics disabled).
`commands/index.ts` builds the root command tree from every `.cmd.ts` module,
so evaluating any one command evaluated all of them. Two of those modules
reached the TypeScript compiler and the code generation pipeline, which
nothing but `composio generate` and `composio run` ever calls:
155.8ms -> 8.0ms commands/run.cmd
63.5ms -> 2.5ms commands/generate
The command tree itself is untouched. `Command.withHandler(input => Effect)`
already runs lazily, so moving these imports inside the handler bodies is
enough. Specs, flags, descriptions, subcommand wiring and `root-help.ts`
introspection all still resolve eagerly, which is why parsing, help rendering
and "did you mean" suggestions cannot shift.
run.cmd.ts was the CLI's only consumer of `import ts from 'typescript'`,
through three source rewrites that `composio run` applies to a user script.
Those move to `run-source-transforms.ts`, which the handler imports
dynamically. The test suite imports them from the new path.
ts.generate.cmd.ts and py.generate.cmd.ts pulled `src/generation/*` at module
scope. Both now resolve it inside the handler, immediately before the first
use.
A rejected import of a module bundled into this binary is an impossible
invariant rather than a recoverable failure, so these use `Effect.promise`
rather than `Effect.tryPromise`. The module registry memoizes each import, so
repeat calls within one run cost nothing.
Behavior is unchanged, checked rather than assumed. Eleven invocations,
covering `--help` at root and for generate, generate ts, generate py, run,
tools and execute, plus `version`, `--version`, an unknown command and an
unknown flag, produce byte-identical stdout, stderr and exit codes before and
after.
Verified with typecheck (src and test), oxlint, validate:boundaries,
validate:skills, and the full package suite: 1326 passed, 1 skipped.
`test/src/cli-main.test.ts` times out in this container and does so identically
on the parent commit (25.6s and 25.2s there, 24.5s and 24.3s here) because it
spawns the CLI from source against a 15s timeout. Not caused by this change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EzaE7oGVgziJ5nRvBhcci2
## Summary
`composio execute` exits 1 whenever a tool response contains the literal
text `<|endoftext|>` or `<|endofprompt|>`. A README about tokenizers is
enough. A 66-byte payload reproduces it.
The tool has already run by then. Side effects happen, stdout stays
empty, no session history entry is written. Someone who sent an email
this way sees a failure and has no response to inspect.
Cause: js-tiktoken's `encode(text, allowedSpecial = [],
disallowedSpecial = "all")`. The CLI passed only the text, so every
special token was disallowed and `encode` threw.
Stacked on #4463. GitHub retargets this to `next` once that merges.
## Changes
`countOutputTokens`, the only `encode` call, now passes `'all'`:
```ts
const countOutputTokens = (json: string): number =>
getExecuteOutputEncoder().encode(json, 'all').length;
```
The encoder is a length gauge for the inline-vs-file decision, so
allowing the literals is the right reading. Each one counts as the
single special token it encodes to, not as the seven tokens its
characters would be. Counts for text without the literals do not change.
The byte pre-filter in #4463 hides this below 10KB by skipping the
encoder. That is cover, not a fix, which is why it stayed out of that
PR.
## Type of change
- [x] Bug fix
- [ ] New feature
- [ ] Refactor/Chore
- [ ] Documentation
- [ ] Breaking change
## How Has This Been Tested?
Bun 1.4.1+4661e494f, Node 24.17.0, pnpm 11.8.0, linux-x64.
1. New test drives `composio execute` through the `cli([...])` harness
with a response holding both literals, sized past the inline threshold
so the count is computed. It reads the stored file back and asserts both
literals survived.
2. Against the unfixed code it fails with `Error: The text contains a
special token that is not allowed: <|endoftext|>`. With the fix it
passes. Suite count goes 90 to 91.
3. `pnpm run typecheck && pnpm run validate:boundaries && pnpm exec
vitest run test/src/commands/tools/tools.execute.cmd.test.ts
test/src/commands/run.cmd.test.ts`
To see it by hand: point any tool at content holding one of the literals
and make the response larger than 10KB.
## Screenshots (if applicable)
Not applicable.
## Checklist
- [x] I have read the Code of Conduct and this PR adheres to it
- [x] I ran linters/tests locally and they passed
- [ ] I updated documentation as needed
- [x] I added tests or explain why not applicable
- [ ] I added a changeset if this change affects published packages
No docs cover the threshold. `@composio/cli` is private, so no
changeset.
## Additional context
o200k defines exactly two special tokens. `encode` builds a regex from
the disallowed set and throws on the first match, before tokenizing.
That guard exists to stop callers from smuggling control tokens into a
model prompt. This code is measuring a string.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01EzaE7oGVgziJ5nRvBhcci2
encode(json, 'all') turns each literal into one special token rather than
counting its characters as text, so say that in the comment and the test name.
`Tiktoken.encode` signs as `encode(text, allowedSpecial = [],
disallowedSpecial = "all")`. The CLI passed only the text, so every special
token was disallowed and the call threw on any response containing the
literal `<|endoftext|>` or `<|endofprompt|>`. A 66-byte payload is enough.
Reading a README that documents a tokenizer hits it.
The throw landed in `prepareExecuteOutput`, after `spinner.stop('Execution
successful')` had already printed. So the tool had run, its side effects had
happened, and the CLI still exited 1 with nothing on stdout and no session
history entry.
Here the encoder is only a length gauge for the inline-versus-file decision,
so those literals are ordinary characters. Passing `allowedSpecial: 'all'`
counts them instead of rejecting the payload. Token counts are unchanged on
text that contains no special tokens.
The preceding commit's byte-length pre-filter hid this below 10KB by
skipping the encoder. Larger responses still reached it, and both call sites
were affected: the threshold check and the `tokenCount` reported for a
stored file. Both now go through one `countOutputTokens` helper.
The regression test drives the real command with a response holding both
literals, sized past the inline threshold so the count is actually computed.
Verified it fails without the fix, with the original error:
Error: The text contains a special token that is not allowed: <|endoftext|>
Verified with typecheck (src and test), oxlint, validate:boundaries, and the
tools.execute and run command suites, now 90 tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EzaE7oGVgziJ5nRvBhcci2
## Summary
`composio --version`: 970ms to 749ms. Peak RSS: 175.8MB to 132.2MB.
Binary: 96MB to 86MB. No new dependencies (one unused one removed), no
API changes.
The compiled binary carried two copies of the TypeScript compiler and
all six tiktoken rank tables. A compiled Bun binary parses everything it
embeds before running any JS, so every command paid for that.
First PR in a stack of five. Review order: this, #4464, #4468, #4469,
#4475.
Bun 1.4.1+4661e494f, linux-x64, best of 7, analytics disabled:
| | before | after |
|---|---|---|
| `composio --version` | 970ms | 749ms |
| peak RSS | 175.8MB | 132.2MB |
| compiled binary | 96MB | 86MB |
| bundle | 31.24MB | 15.84MB |
## Changes
1. `run.cmd.ts` imported `ts` from ts-morph, which bundles its own
TypeScript. It now uses the `typescript` package the generation code
already pulls in. One compiler instead of two, 8.5MB each.
2. js-tiktoken's main entry inlines six rank tables (5.3MB). The CLI
only uses o200k. Switched to `js-tiktoken/lite` with that one table.
Token ids are identical.
3. `prepareExecuteOutput` built the rank table on every successful
execute just to compare against a 10,000 token threshold. A token covers
at least one byte, so a payload under 10,000 bytes cannot exceed 10,000
tokens. It checks bytes first and only builds the tokenizer past that.
It also checks the invocation origin before building it, since `composio
run` always prints inline, and a stored response is encoded once rather
than once for the threshold and again for `tokenCount`.
4. `ToolsExecutorLive` resolved a client via `clientSingleton.get()`
(disk reads) and then discarded it, since every caller passes one in.
Resolved lazily now.
5. `ts-morph` is removed from the CLI's dependencies. Nothing imports it
after change 1, and leaving it listed made it easy to bring its
TypeScript copy back.
What changes in behavior:
- Responses under 10KB no longer reach `Tiktoken.encode()`, so the
`<|endoftext|>` crash stops happening for them. Larger responses still
hit it. #4464 is the real fix.
- Executes started by `composio run` no longer build the tokenizer for
responses over 10KB. Their output was always printed inline, so the
count was discarded.
- TypeScript 6.0.2 (ts-morph's copy) becomes 6.0.3.
- Under `COMPOSIO_LOG_LEVEL=Debug`, ProjectContext's "resolved from"
lines no longer print on the remote execute path.
On change 3: an earlier version of this description said it saved ~500ms
per execute, measured in isolation under a different Bun. Inside the
compiled binary, `new Tiktoken(o200k)` costs ~390ms and `encode()` of
7.5KB about 4ms. Measured end to end on a real
`HACKERNEWS_GET_ITEM_WITH_ID` execute with `COMPOSIO_PERF_DEBUG=1`, the
time from `execute.tool_call.end` to exit drops from ~300ms to ~15ms for
responses under 10KB, and stays ~350ms above it.
## Type of change
- [ ] Bug fix
- [ ] New feature
- [x] Refactor/Chore
- [ ] Documentation
- [ ] Breaking change
It removes a crash, but by accident, so it is not marked as a bug fix.
## How Has This Been Tested?
Bun 1.4.1+4661e494f, Node 24.17.0, pnpm 11.8.0, linux-x64.
1. `cd ts/packages/cli && pnpm run typecheck && pnpm run
validate:boundaries`
2. `pnpm exec vitest run
test/src/commands/tools/tools.execute.cmd.test.ts
test/src/commands/run.cmd.test.ts`: 90 passed. A new case covers a ~18KB
response that encodes to ~4k tokens, past the byte check but under the
threshold, and asserts it stays inline.
3. `pnpm build:binary && time ./dist/composio --version`
Checked but not committed: the three parse helpers give identical output
under both compilers across 20 sources (TSX, decorators, `using`,
`satisfies`, import attributes, unicode). Lite tiktoken gives identical
token id streams on 8 samples including CJK, RTL, emoji and control
characters. Max tokens per byte was 0.846, under the 1.0 the byte check
needs.
## Screenshots (if applicable)
Not applicable.
## Checklist
- [x] I have read the Code of Conduct and this PR adheres to it
- [x] I ran linters/tests locally and they passed
- [ ] I updated documentation as needed
- [x] I added tests or explain why not applicable
- [ ] I added a changeset if this change affects published packages
No docs describe the bundle contents or the token threshold. The byte
pre-filter boundary has a test. The compiler and tokenizer comparisons
above are still not committed and should become a suite. `@composio/cli`
is `private: true`, so no changeset.
## Additional context
Bun 1.4.2 gives no gain over 1.4.1 (three rounds of best of 7:
723/732/756ms vs 744/713/747ms). Keep the pin.
Not touched: ~1.1s of execute preflight (5 to 7 serial round trips;
`project/resolve` has no cache and can fire three times), the
two-round-trip session create plus execute, and `--skip-checks`, which
currently skips nothing measurable.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01EzaE7oGVgziJ5nRvBhcci2
Check the invocation origin before building the tokenizer, since composio run
always prints inline, and pass the token count from the threshold check into
the stored-output summary instead of encoding the payload twice.
Add a test for a response past the byte pre-filter but under the token
threshold, and drop the unused ts-morph dependency.
A compiled Bun binary parses its whole embedded module graph before the
first line of JS runs, so bundled-but-unused code is paid for on every
invocation. Verified: a binary that bundles everything but evaluates only
console.log still costs ~235ms, against 15ms for a hello-world build.
Most of this bundle was code `composio execute` never reaches.
Measured on the pinned toolchain (Bun 1.4.1+4661e494f, linux-x64,
best of 7, analytics disabled):
composio --version 970ms -> 749ms (-221ms)
peak RSS 175.8MB -> 132.2MB (-43.6MB)
compiled binary 96MB -> 86MB
bundle 31.24MB -> 15.84MB
Four changes.
run.cmd.ts imported `ts` from ts-morph, which vendors its own copy of the
TypeScript compiler, so the binary carried two of them. The file uses only
createSourceFile, forEachChild, ScriptTarget, ScriptKind and five isX
guards, all available in the typescript copy that
src/generation/typescript/* already pulls in. Sharing one compiler also
means commands/generate and commands/run.cmd no longer evaluate a compiler
each.
js-tiktoken's main entry statically inlines all six BPE rank tables (gpt2,
r50k, p50k, p50k_edit, cl100k, o200k) as string literals; the CLI only ever
uses o200k, via encodingForModel('gpt-4o'). The lite build with that single
table produces identical token-id streams.
prepareExecuteOutput built the o200k rank table on every successful
execute, purely to compare the response against a 10k-token threshold.
Constructing it measured ~390ms in a compiled binary on the pinned
toolchain (390.1, 367.1, 418.6ms across three runs), against ~4ms to
encode a 7.5KB payload once the table exists. A BPE token always covers at
least one UTF-8 byte, so a payload of at most THRESHOLD bytes can never
exceed THRESHOLD tokens; checking byte length first reaches the same
decision without the tokenizer.
That ~390ms is construction cost measured in isolation, not an end-to-end
delta on a real `composio execute`. No credentialed run was available to
measure the whole command before and after, so treat it as the size of the
work removed from the success path rather than a verified wall-clock
saving. `COMPOSIO_PERF_DEBUG=1` reports the gap between
`execute.tool_call.end` and exit for anyone able to run it for real.
ToolsExecutorLive resolved a client through clientSingleton.get()
unconditionally, walking the project context off disk, then discarded it
because every caller on the remote-execute path passes one in.
Three behavioral deltas, none of them the tokenization result or the
inline/file decision:
1. Tiktoken.encode() throws on the literals <|endoftext|> and
<|endofprompt|> appearing anywhere in the response, at any size (a
66-byte payload reproduces it). That throw landed after
"Execution successful" had printed, so the tool ran and the CLI still
exited 1 with nothing on stdout. Responses at or under 10KB no longer
reach encode(), so they now succeed. Larger responses still hit it;
the real fix is passing allowedSpecial 'all' and is not in this commit.
2. TypeScript 6.0.2 (ts-morph's vendored copy) to 6.0.3. Differential
tested: the three real parse helpers over 20 sources covering TSX,
decorators, `using`, `satisfies`, import attributes, optional-chained
calls and unicode gave identical output on all 60 comparisons.
3. Under COMPOSIO_LOG_LEVEL=Debug, ProjectContext's "resolved from ..."
debug lines no longer appear on the remote-execute path. The local-tool
path still calls get() and is unchanged.
Verified with typecheck:src, oxlint, validate:boundaries, the
tools.execute and run command suites (89 tests), and differential tests of
both tokenizers and both TypeScript versions.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EzaE7oGVgziJ5nRvBhcci2
## Summary
Automated refresh of the toolkit slugs the CLI knows without asking
the API, generated by
`ts/packages/cli/scripts/generate-toolkit-slugs.ts`.
Toolkits added since the last refresh currently cost users one
toolkit-list fetch (~2 s) the first time they run one of that
toolkit's tools. Merging this makes them free.
The generator refuses to write a list that is short, malformed, or
missing staple toolkits, so a bad fetch opens no PR at all.
This PR:
- closes [PRDE-1613](https://linear.app/composio/issue/PRDE-1613)
- maps only 404/400 from `tools.retrieve` to
`ComposioToolNotFoundError`; every other failure (401 invalid API key,
5xx, network) now raises the new `ComposioToolFetchError` with the
client error kept as `cause`
- `tools.get(userId, slug)` and `tools.execute` inherit the corrected
mapping since they call `getRawComposioToolBySlug`
- fixes `toolkits.get(slug)`, whose 404/400 check compared against the
OpenAI `APIError` class instead of the Composio client one, so
`ComposioToolkitNotFoundError` never fired
- Python parity: `get_raw_composio_tool_by_slug` raises
`ToolNotFoundError` (now a `NotFoundError` subclass) on 404/400 and
re-raises any other `composio_client` error unchanged
- adds unit tests on both sides for 404, 400, 401 and non-API failures;
verified live against the API with an invalid key on both SDKs
## Context
An unauthenticated call to `tools.getRawComposioToolBySlug` returned
`error.name === "ComposioToolNotFoundError"` while `error.cause.status`
was 401. The catch block wrapped every error except cancellation as
not-found, which predates the `@composio/client@beta` swap. Intended to
be back-ported to `main` after merging to `next`.
https://claude.ai/code/session_017HtbhwMAKcfebo8HyXWa5s
getRawComposioToolBySlug relabelled every client error, including an
invalid API key (401), as ComposioToolNotFoundError. Map only 404/400 to
not-found and wrap the rest in a new ComposioToolFetchError that keeps
the client error as cause. Toolkits.getToolkitBySlug compared against
the OpenAI APIError class, so its not-found branch never fired; import
the Composio client class instead. Python mirrors the mapping: an
unknown slug raises ToolNotFoundError (now a NotFoundError), anything
else propagates the composio_client error unchanged.
PRDE-1613
Claude-Session: https://claude.ai/code/session_017HtbhwMAKcfebo8HyXWa5s
## Summary
Route documentation contact links to the Composio contact form with both
`utm_source=docs` and page-specific CTA placement tracking.
## Changes
- Replace Calendly links on both rate-limit pages with
`https://composio.dev/contact?utm_source=docs&cta_placement=docs-rate-limits`.
- Update the data-retention sales link to use
`utm_source=docs&cta_placement=docs-data-retention`.
- Replace the premium-tools billing email link with the contact form
using `utm_source=docs&cta_placement=docs-pro-tools`.
## Type of change
- [x] Documentation
## How Has This Been Tested?
Verified that all four Markdown contact links contain the expected
source and page-specific placement parameters with Python URL parsing
assertions. `git diff --check` passed. The docs build and test suite
were not run; docs dependencies are not installed in this checkout.
## Checklist
- [x] I updated documentation as needed
- [x] I added tests or explain why not applicable
No new tests or changeset are needed for these four documentation URL
replacements.
## Summary
Explain the motivation and context for this change. Link to any related
issues.
Fixes #
## Changes
-
-
## Type of change
- [ ] Bug fix
- [ ] New feature
- [ ] Refactor/Chore
- [ ] Documentation
- [ ] Breaking change
## How Has This Been Tested?
Describe the tests you ran and instructions so reviewers can reproduce.
Include any relevant config/versions.
## Screenshots (if applicable)
## Checklist
- [ ] I have read the Code of Conduct and this PR adheres to it
- [ ] I ran linters/tests locally and they passed
- [ ] I updated documentation as needed
- [ ] I added tests or explain why not applicable
- [ ] I added a changeset if this change affects published packages
## Additional context
This PR:
- Add `.github/workflows/agent-substrate.yml` running `pnpm
validate:agent-skills` and `pnpm validate:skill-routing` on every push
and pull request; both validators previously ran in no CI workflow
- No path filters on the trigger: the stale-guidance walk scans every
text file in the repo, so any change can affect the result (PR runs
restore caches but only `next` pushes save them, per the
`setup-node-pnpm-bun` guidance)
- Skip `vendor/` directories in the `validate:agent-skills`
stale-guidance walk, which was failing on read-only third-party
snapshots mentioning other tools' rule conventions
- Extend the validator's command scan to `CONTRIBUTING.md` (with a `pnpm
dlx` exemption), so its documented commands are checked against
`package.json`, `python/Makefile`, and `python/noxfile.py` like the rest
of the guidance
- Point the routing-test header, root `AGENTS.md`, and
`skill-maintenance` reference docs at the new workflow, and add a
"Working with AI Coding Agents" section to `CONTRIBUTING.md` covering
the inherited agent setup, the two checks, and the routing-probe
requirement for skill edits
## Context
These two validators are the only checks keeping repo-level agent
guidance honest: command names mentioned in guidance are verified
against `package.json`, `python/Makefile`, and `python/noxfile.py`, and
routing probes assert each skill stays the unique top match for its
representative task. Until now nothing enforced either one, and the
stale-guidance walk was already red on vendored trees — a failure no
guidance owner could fix, which trains people to ignore the check. This
makes both checks blocking everywhere they can bite.
## Verification
- `pnpm validate:agent-skills` — 19 skills, green, now including
`CONTRIBUTING.md` commands
- `pnpm validate:skill-routing` — 19 probes over 19 skills, green
- Workflow YAML parsed; oxlint and prettier clean on touched files
- `Agent Substrate` workflow ran green on this PR (42s) before the
trigger change and re-runs on every push
## Summary
Autonomous coding agents assumed Composio signup required a human,
leaving integrations without credentials or live verification. Add a
prominent guide for the supported `composio login --agent` flow when no
human is available.
Addresses Gauge action `cmtvu1rwe00040ip8mxhszmj1`. Reviewed the
metadata, insights, logs, and diffs for [evidence run
1](https://agents.withgauge.com/composio-aclx/runs/cmtvrts3n001201ea7jwbhu4q)
and [evidence run
2](https://agents.withgauge.com/composio-aclx/runs/cmtvrts3n001401eak0rwndar).
Both assumed signup required human interaction; the second attempted
disposable-email signup before switching to mocks.
## Changes
- Document unattended login, readiness checks, project API-key
extraction, credential handling, constraints, and the human login
fallback.
- Include a live Hacker News tool call and require separate verification
of the requested integration, including provider authorization and
confirmation of write results.
- Link the guide from the agent setup sidebar, quickstart, CLI docs, API
authentication reference, and `llms.txt`.
## Type of change
- [x] Documentation
## How Has This Been Tested?
From `docs/`:
- `bun run test`: 557 passed, 0 failed. Run with loopback access for the
analytics test's local server.
- `bun run lint:links`: 0 errors.
- `bun run lint`: passed with existing warnings.
- `bun run postinstall`: regenerated MDX collections successfully.
All five shell snippets pass `bash -n`. The key-extraction snippet
accepts a valid local fixture and rejects missing, blank, and non-string
keys without printing credentials. `git diff --check` passes. No live
account was provisioned or tool executed for this documentation change.
## Checklist
- [x] I ran linters/tests locally and they passed
- [x] I updated documentation as needed
- [x] I added tests or explain why not applicable: existing docs checks
and focused snippet validation cover this documentation-only change.
- [x] I added a changeset if this change affects published packages: not
applicable; no published package changes.
Automated knowledge-base refresh for `ComposioHQ/support-knowledge`.
- Source commit: `5eac683455ff252a7a3b62f33ab6566445009b52` (unchanged;
rebuilt a stale semantic artifact)
- Regenerated public KB pages and search records
- Reused unchanged vectors and rebuilt the checked semantic artifact
- Ran KB freshness and semantic-artifact verification
Co-authored-by: sohambasu963 <80603154+sohambasu963@users.noreply.github.com>
## Summary
Editing an indexed docs page makes the checked-in semantic artifact
stale, causing `Docs - Tests` to fail until a credentialed rebuild
commits new embeddings. Make freshness advisory in PR checks so docs
changes can proceed while the existing refresh workflows rebuild the
artifact.
## Changes
- Add `--allow-stale` to the PR semantic check. Stale artifacts emit a
warning; missing artifacts and integrity failures still fail the check.
- Validate artifact integrity before reporting freshness mismatches, so
stale content cannot hide corrupt vectors or invalid records.
- Keep rebuild workflows and runtime validation strict. Search uses its
existing keyword fallback until a fresh artifact is available.
- Document the contributor workflow and add regression coverage for
freshness, corruption, and CI wiring.
## Type of change
- [x] Bug fix
- [x] Documentation
## How Has This Been Tested?
Run from `docs/` with Bun 1.4.2:
- `bun test tests/static/kb-semantic-artifact.test.ts
tests/static/kb-update-workflow.test.ts
tests/static/knowledge-search.test.ts`: 72 passed.
- `bun run test`: 567 passed; one analytics test could not open its
local HTTP server in the sandbox. `bun test
tests/static/kb-query-analytics.test.ts` passed all 14 tests when rerun
with permission to listen.
- `bun run types:check`: passed.
- `bun run lint`: passed with existing warnings.
- `bun run check:kb-semantic` and `bun run check:kb-semantic
--allow-stale`: passed for the current artifact.
- CLI smoke checks with temporary artifact changes confirmed that stale
content exits 1 in strict mode and exits 0 with a GitHub warning in
advisory mode. A stale artifact with corrupt vector data still exits
nonzero. The original artifact was restored.
## Checklist
- [x] Ran linters and tests locally, with results recorded above
- [x] Updated documentation
- [x] Added regression tests
- [x] Changeset not required: docs-site and CI changes only
## Additional context
Merging stale embeddings temporarily reduces search to keyword retrieval
until a refresh lands. Embeddings are still generated by the existing
background workflows, without adding API calls to ordinary PR checks or
generating the corpus during search requests.
## Summary
Auto-generated Python SDK reference docs from `python/composio/`.
Regenerates pages at `docs/content/reference/sdk-reference/python/` to
reflect changes in the Python package's public API (new methods, updated
signatures, changed types).
`PackageInstall` disappeared from agent-readable Markdown, removing
required package commands from the quickstart, provider guides, and
search records. Preserve Node and Python install commands,
package-choice comments, coding-agent setup prompts and client links,
and video captions. Share the UI's package-manager definitions and setup
content with the converter.
Addresses [DEVREL-35](https://linear.app/composio/issue/DEVREL-35). The
component audit records existing coverage and remaining visual/catalog
parity work for DEVREL-32 and DEVREL-38.
Validation from `docs/`:
- Static suite: 553 passed.
- Typecheck, lint, and production build passed. Lint retains existing
warnings.
- Built-server HTTP regression: quickstart, Anthropic provider, client
setup, and `/llms-full.txt` retain the expected commands and prompts.
- Corpus regression covers every authored PackageInstall instance in
page text and search records, plus raw and processed attribute forms.
The semantic artifact content hash is stale because indexed text
changes. CI fails at that check. The automatic rebuild skipped its
membership gate on both the initial run and one retry, even though the
current PR API reports MEMBER. A maintainer needs to run the authorized
artifact rebuild before merge. No OPENAI_API_KEY is available locally;
no embeddings were edited by hand and no search index was published
locally.
## Summary
The docs homepage used the article OG template with “Welcome” as its
headline. Its page-level metadata also dropped the site name. Add a
dedicated homepage card with “Build and operate AI agents,” SDK/CLI/MCP
entry points, and a “Start building” prompt.
## Changes
- Route the docs homepage and default social image to the new 1200×630
card. Preserve article cards.
- Use a descriptive homepage title and shorter description, and set
explicit Open Graph and Twitter metadata.
- Cover homepage image routing and rendered metadata after the root
redirect with regression tests.
## Type of change
- [x] Bug fix
- [x] Documentation
## How Has This Been Tested?
Commands run from `docs/`:
- `bun test tests/static/og-image.test.ts`: 2 passed, including article
and homepage PNG rendering.
- `TEST_BASE_URL=http://localhost:3123 bun test
tests/integration/rendering.test.ts --test-name-pattern 'docs homepage
social metadata' --timeout 60000`: passed against the local dev server.
- `bun run types:check`: passed.
- `bun run lint`: passed with warnings in unrelated files.
- Rendered and visually inspected the homepage PNG at 1200×630.
- `git diff --check`: passed.
Image tests were rerun after updating to current `next`. Remote CI and
production deployment are pending.
## Checklist
- [x] Ran local checks as described above
- [x] Updated homepage metadata
- [x] Added regression tests
- [x] No changeset required: docs-only change
Replace the exhaustive `/llms.txt` dump with a short routing map for
product selection, installation, authentication, sessions, execution,
and troubleshooting. Keep the full catalog at `/llms-index.txt`, using
named links and descriptions, and put that catalog and other long-tail
resources under Optional. Current REST v3.1 and legacy v3.0 remain
explicitly separated.
Fixes [DEVREL-34](https://linear.app/composio/issue/DEVREL-34). The
format follows the descriptive-link and Optional conventions in the
[llms.txt proposal](https://llmstxt.org/).
Validation: 551 static tests passed, including bounded routing-map
coverage and route resolution. Typecheck, lint, and link validation
passed. Existing exhaustive-catalog coverage now tests
`/llms-index.txt`; the HTTP version-grouping test follows the new
catalog route. Existing lint warnings remain.