The agents list verbs carried their own flat has_more / page_token fields on the
envelope meta, while the rest of the CLI had moved to the structured
Meta.Pagination block. That left two pagination vocabularies in one envelope.
Drop HasMore / PageToken from output.Meta and have listMetaPage populate
PaginationMeta instead. The agents verbs are a single-page cursor API — one call
fetches one page and returns a cursor — so Pages is fixed at 1 and Complete
reads as "this page is the last one". Count is left unset so Items is the single
source for the record count, avoiding the same number twice.
Non-paginated paths (the provider listing and catalog enumeration) still go
through listMeta and keep meta.count.
Update the affected assertions and the three skill reference docs, including the
consumption guidance that still pointed at meta.count on the paginated leaves.
Address the two blockers the pre-publication OPSEC review raised against this
subtree, and drop the in-repo demo provider that is not meant to ship.
Remove the example provider:
- delete agents/example (the offline demo backend and onboarding template)
- migrate the cmd-layer tests onto a new catalog-type test fake (fakecat,
registered beside the existing instance fakes), so the catalog-specific
paths stay covered — unknown agent id, per-agent capability differences,
declared send params, per-operation brand scoping
- delete the example provider skill doc; retarget the remaining docs at base,
using output captured from real `agents list` / `agents card` runs, and
relabel the samples that can no longer be verified as structural examples
Publication norms:
- translate every user-facing string in the subtree to English (command help,
flag usage, error messages, hints, meta.next labels), plus test messages
and fixtures. Chinese now remains only in skills/**.md, the sanctioned
convention, and in the test that asserts against those docs
- drop the dangling internal design-doc section references that a public
reader cannot resolve; public-standard RFC citations are kept
- re-sync the quoted CLI output in the skill docs with the new messages
Behavior changes carried in this commit:
- run the capability gates before the --dry-run branch, so a preview is no
longer a capability bypass: --answer / --file against an agent declaring
input_required=false / file_input=false now report unsupported_capability
under dry-run too, matching what a real send would answer. The --file
confirmation gate stays after dry-run, which uploads nothing
- raise the meta.next watch window from 30s to 90s. A typical backend task
runs ~60s, so a 30s window made a caller re-issue --watch two or three
times, and each new process restarted the poll backoff with no gap between
calls — enough to self-hammer the backend into a rate limit
Also trim the longest narrative comment blocks in the subtree.
Replace the provider-level flat RequiredScopes set with per-identity
declaration (IdentitySpec.Scopes, never serialized into the card): user
and bot each declare their own set, and the preflight resolves the
required scopes from the RESOLVED identity before reading the granted
list, so an identity with no declared scopes skips the check entirely --
for bot that also skips the app-version TenantScopes fetch.
Register now rejects duplicate identity types (ambiguous once scopes
attach to identities) and empty/duplicate scopes within one identity;
agenttest conformance mirrors the same rules. The check itself is
unchanged: all-or-nothing per identity, missing_scope error shape and
identity-specific remediation hints as before, --dry-run untouched.
New fixture fakesplit (user declares one scope, bot none) pins the
per-identity resolution and the bot fetch skip. Skill docs updated to
the per-identity wording.
Normalize the first valid log ID from the response body or transport headers before API error classification. Cover nested, whitespace, invalid-type, and header fallback cases.
Keep static agent cards discoverable while preventing unsupported identities from reaching dynamic provider operations. Apply the same validation to online agent enumeration and dry runs.
1. Update Base Agent endpoints, scopes, params, pagination, and response decoding.
2. Map task schema v1 outputs into messages, artifacts, and input-required details.
3. Preserve raw structured data and output metadata across polling and rendering.
4. Add unit tests, dry-run E2E coverage, and provider documentation.
Replace the decision_id/--option HITL contract with the question-group
answer scheme:
- contract: InputRequired becomes {label, description, questions[]}
(a single question is a length-1 group); Option gains description;
decision_id/input_type/submitted fields are removed
- answering: --answer <qid>=<option_id> | <qid>.text=<text> is the only
answer channel (--decision-id/--option removed); --text is always the
message-level remark; offline collect-all key-grammar guards plus an
input_required capability gate
- keys: minted at group-creation time (MintQuestionIDs /
DeriveGroupSuffix), per-group-unique for stale-retry protection;
central normalization with size caps and whole-group degradation on
non-conforming keys (JSON _notice.provider_defect); KeyCharsetRE is
the single charset source and safeNextID now requires an alphanumeric
first character (rejects flag-lookalike ids)
- meta.next: per-question --answer templates with a relay-first label
- example/planner: three-question reference group; strict collect-all
validation (Reason enum + question spec), atomic acceptance with a
resolved_answers echo, context binding, no sibling-task fork on bare
--text; Register enforces InputRequired => brand-covering CancelTask
- errs: ValidationError gains the resolved_answers extension field
- skills/lark-agents 1.3.0: relay-first rules, answer grammar, recovery
playbook, untrusted question-text discipline
The set of agents and the capabilities each exposes can be scoped to the
resolved account environment, so discovery and gating reflect only what the
current account can actually use.
SPI (internal/agents):
- AgentSpec and Op gain an optional scope field (empty = unrestricted);
Register fail-fasts on an unknown scope value and on a scope declared on an
unwired operation.
- DeriveCapabilities and BuildCard take the resolved environment: an
operation-backed capability is exposed only when wired AND in scope; the card
echoes the environment it was rendered for.
- Provider.ListCatalog filters the catalog to the in-scope agents.
Command layer (cmd/agents):
- The resolved environment defaults offline (nil-safe) so the gates hold before
config init. Every verb path gates offline before the client is built: the
whole-agent gate fires before the per-verb capability nil-gate, so an
out-of-scope agent uniformly returns a single dedicated validation error
(exit 2) for every verb instead of a misleading "unsupported" on an
incidentally-unwired one; the per-operation gate follows the nil-gate. List
filtering mirrors the same rule.
The example provider demonstrates the per-capability scope end to end; skill
docs updated accordingly.
Adds Feishu-OpenAPI-style page_token/page_size pagination to the three list
operations (`agents task list`, `agents context list`, instance
`agents list <scheme>`), consistent with the shortcuts/contact & calendar
convention.
SPI (internal/agents):
- New PageParams{Token,Size} / PageInfo{NextToken,HasMore}.
- ListTasks/ListContexts hooks and the Provider.ListAgents field gain a
PageParams arg and a PageInfo return; the provider owns cross-page ordering
(contract: most-recent-first) and maps the opaque cursor to its backend.
Command layer (cmd/agents):
- --page-size (default 20, range 1-100, validated client-side in RunE) and
--page-token on the three list leaves. output.Meta gains has_more +
page_token (both omitempty); listMetaPage emits them and, when a next page
exists, a ready-made "下一页" meta.next command so an AI pages by running the
suggested command instead of threading a cursor. The next-page cursor is
safeNextID-whitelisted before it is interpolated into that command (it still
rides meta.page_token as data if it fails), and ref/scheme/context-id are
each whitelisted too. The catalog list path stays offline/unpaged.
- Per-page CLI re-sort removed: ordering is now purely the provider's
responsibility, so a within-page re-sort could only make the concatenation
across pages inconsistent.
example provider paginates its in-memory store via an opaque offset cursor,
most-recent-first (Seq desc). Skill docs (task/context/list) document the flags,
has_more/page_token, and the meta.next paging idiom.
Renames the provider-neutral agent surface from singular to plural across
every layer, for internal/external consistency:
- CLI command `lark-cli agent` → `lark-cli agents` (cobra Use, the
NewCmdAgents constructor, root registration, and the Agent-tooling help
group key so it still groups correctly).
- Go packages / dirs: agent/→agents/, cmd/agent/→cmd/agents/,
internal/agent/→internal/agents/; every package declaration, import path,
and the iagent→iagents import alias.
- Skill: skills/lark-agent/→skills/lark-agents/, reference files
lark-agent-*.md→lark-agents-*.md, SKILL.md frontmatter name/cliHelp, and
every documented command reference.
Domain nouns deliberately stay singular — they model one agent, not a
collection: the AgentSpec/AgentCard/AgentTask/AgentSummary types, the
agent_id / agent_ref / agent_ref_format / agent_id_source wire keys, and the
A2A message Role "agent". Three adversarial review passes confirmed no
wire-key or type corruption and no missed references.
Code (all findings verified by re-run before fixing):
- send --file: local gate before any capability/confirmation gate —
relative-within-CWD, must exist, not a directory; violations collected
in one pass (dry-run included)
- --param scalars: canonicalize accepted variants (TRUE/1/+5/04 ->
true/5/4) so both channels produce one wire form; integer overflow now
reports a range error, not a type error; non-object JSON no longer
leaks Go type text
- unknown-param suggestions: nearest-first by edit distance (<=2),
cross-verb hits no longer mix verb names into suggestions
- meta.next: no self-loop on terminal task get (caller-aware); --as is
carried only when the caller passed it explicitly, matching the
shortcut-family convention of never pinning identity
- pretty task view: request/reply lines (result was invisible before)
- context delete gate: self-contained Chinese irreversibility message
- teaching hints: --timeout/--task-id now carry actionable fixes;
example conflict message drops internal jargon
- list-class envelopes: empty lists omit meta entirely (no '{}' shape)
- example: artifacts carry name/mime before download (csv/xlsx);
planner now covered by conformance
Docs: skill family synced to verified binary behavior — bot preflight
truth, incremental-auth hint semantics, 8-verb
--operation vocabulary, real-output sample backfills (3 agents,
created_at, submitted omitempty), redundancy converged to single
authority sections, provider file slimmed to runtime-AI-only content.
Object parameters land on the declaration model: CardParam gains Fields (one
nesting level, scalar leaves only — an object itself declares nothing but its
members; requiredness, enums, defaults and ranges all live on the leaves) and
Register recursively fail-fasts six new object rules.
Transport is dual-channel with one canonical form: dotted paths are the
primary channel (--param filter.region=east — no shell quoting, AI error rate
stays at scalar level, and meta.next can carry leaves literally) and a JSON
value is the fallback (--param filter='{"region":"east"}', numbers decoded via
json.Number so literals survive). Both channels validate leaf-by-leaf with the
same teaching errors (dotted-path violation names, enum/range sets, unknown
members listing the field set) and normalize into flat dotted keys in
rt.Params() — a provider never sees which channel the caller used. Mixing
channels for one object is rejected; leaf defaults backfill on both.
Consumption: ParamObject[T] assembles an object's leaves into a typed struct;
BindParams supports nested tagged structs; agenttest.CheckParamsBinding
recurses so declaration/struct drift on leaves still dies in CI.
NoCarry opts a parameter out of the meta.next carry: values never ride the
chain literally (a required NoCarry param degrades to a placeholder so the
caller supplies a fresh value). It is the declared escape hatch for per-call
parameters (trace tags) where the carry rule's same-resource continuity
assumption does not hold. Object leaves otherwise carry as ordinary scalars
under the unchanged three-way rule.
The example reporter declares a render object (enum + boolean + defaults)
consumed through a nested binding struct — with defaults the historical reply
stays byte-identical. Skill docs updated (card fields/no_carry semantics, send
dual-channel rules, example walkthrough).
Implements the v5 params design: business parameters become first-class,
per-operation declarations that every agent verb can carry, validate, and
chain — replacing the send-only, agent-global parameter table.
SPI (internal/agent):
- AgentSpec's eight hook funcs become eight Op units (Op[H]{Params, Handler}):
a parameter physically cannot be declared on an unimplemented operation, and
capability derivation enumerates one table (Ops()) shared by every consumer.
Handler signatures are unchanged; migration is wrapping `Send: f` into
`Send: SendOp{Handler: f}`.
- CardParam gains real validation semantics: Type (string|integer|number|
boolean, empty normalizes to string), Enum (string+integer), Default
(backfilled), Min/Max (numeric ranges). Register fail-fasts on twelve
declaration mistakes (charset, duplicates, enum/default/range conflicts,
params on unwired ops, ListParams without ListAgents).
- Runtime gains Params() with a hard contract: required keys present and
non-empty, defaults backfilled, every value validated — before any handler
runs. Typed consumption via BindParams[T] (param struct tags) plus
ParamInt/ParamBool; agenttest.CheckParamsBinding locks declaration↔struct
drift in CI. SendInput.Params and CardInfo.Parameters are removed.
Command layer (cmd/agent):
- --param key=value on all eight verb leaves plus `agent list <scheme>`
(Provider.ListParams; discovered via providers[].list_parameters since the
caller holds no agent_ref at list time). `task get --artifact` validates
strictly against the artifact_download declaration.
- Collect-all validation: every violation reported in one typed error
(errs.InvalidParam gains an optional Spec field embedding the full
declaration), so a caller can fix all mistakes in one round without a
discovery trip. Empty values count as "not provided" (defaults still
backfill; required still rejects). send gains an explicit mode discriminator
(start/continue/answer) formalizing the existing guards.
- meta.next carries params by the three-way rule (whitelisted values ride
literally; whitelist failures degrade required params to placeholders;
target-verb requireds the caller never gave are added as placeholders), and
terminal tasks emit a ready-made download command per artifact.
- Lean card: parameter details move behind `agent card <ref> --operation
<verb|all>` (returns supported/command-template/parameters); the default
card carries a has_parameters cue. Instance providers are labeled
parameters_source:"template".
The example provider migrates to Op units and reporter declares two demo
params (enum+default, integer+range+default) consumed via BindParams — the
copy-start template now exercises the whole declaration surface offline.
Skill docs (SKILL.md + six references) updated to the new discovery and error
contracts.
An adversarial review pass (16 findings, all verified) is folded in — notably:
empty-valued optional params no longer bypass validation and default backfill
(was an internal-error/exit-5 path violating the rt.Params() contract);
invalid values no longer double-report as missing; BindParams returns a typed
error on unexported tagged fields; ValidateValue rejects non-finite numbers.
Restores the structured HITL decision that the initial cut flattened to a prompt
+ free text, so multi-endpoint arbitration is expressible. Grounded in A2A: the
decision rides a DataPart and the answer reuses `agent send` (A2A's single
SendMessage), not new top-level protocol fields or a new verb.
Read side:
- InputRequired gains DecisionID, InputType (single_select|multi_select|text),
Options[]{OptionID,Label}, Submitted (+SubmittedOptionID); Options was []string.
- `task get` pretty view renders the decision (prompt, decision_id, options,
submitted); every agent-controlled field is ANSI-stripped.
Write side:
- SendInput gains DecisionID + OptionIDs. `agent send` gains --decision-id and
--option (repeatable); --text becomes optional when answering by option.
Guards: --option needs --decision-id; --decision-id needs --context-id/--task-id.
- meta.next for an input_required task carrying a decision points at the
structured answer command (decision_id whitelisted via safeNextID; falls back
to --text otherwise).
- Arbitration is server-side: a provider's Send returns conflict on an
already-answered decision; the CLI only reads/echoes submitted.
Reference provider: a new example:planner agent demonstrates the loop end to end
— the first send opens the decision, --decision-id/--option completes it and
records the winning option, and a re-answer returns conflict.
Skill docs (SKILL, send, card, example) updated; card capability count corrected
7 -> 9.
The typed upload helper CallUpload[T] had no caller anywhere — only its
JSON counterpart Call[T] was exercised (by TestCmdRuntime_CallAPI_UnwrapsData)
— so the incremental dead-code gate flagged it as newly unreachable.
Add TestCmdRuntime_CallUpload_PropagatesError, mirroring the Call[T]
coverage: it drives CallUpload through the unsafe --file reject path and
asserts the validation error propagates and T stays zero. This keeps the
provider-facing typed pair (Call[T] / CallUpload[T]) symmetric — both entry
points are now tested rather than only the JSON one — and makes the helper
reachable so the gate passes.
The instance/ListAgents branch of `agent list <scheme>` makes a real online
call but skipped the two gates every other online verb runs via resolveSpec +
preflightScopesForRef: the user|bot identity whitelist and the all-or-nothing
scope preflight. A future network-backed provider would get a worse contract on
list than on send/task/context — a scope-lacking user hit a raw platform error
instead of a clean missing_scope (exit 3), and an out-of-whitelist identity was
never rejected.
- preflightScopesForRef is split into a scheme-keyed core
(preflightScopesForScheme) plus a thin ref wrapper; ref-addressed callers are
unchanged.
- The online list branch now calls f.CheckIdentity and preflightScopesForScheme
(the full RequiredScopes, same all-or-nothing rule as the other verbs) before
ListAgents. Enumeration is treated as a real API verb, not a special case.
- `agent list` registers --as so the enumeration identity can be chosen (the
offline no-scheme provider listing ignores it).
Catalog providers (offline enumeration) are unaffected. Tests cover the scope
preflight (missing_scope, ListAgents not called), the identity whitelist, and
the --as flag registration.
Two review follow-ups on the agent command tree.
Artifact download (data integrity): fetchArtifactURL now reads one byte past
maxArtifactBytes and refuses a body over the cap with a typed error, instead of
letting io.LimitReader silently truncate a >256 MiB artifact to a corrupt,
partial file that reported success (exit 0). The LimitEnforced test is inverted
to assert the rejection.
Capabilities: the single multi_turn card bit is replaced by three independent
capabilities — context_list / context_get / context_delete — each derived from
its own wired hook (ListContexts / GetContext / DeleteContext). One bit could
not honestly represent three separately-deliverable verbs: a provider wiring
ListContexts but not DeleteContext advertised multi_turn=true while `context
delete` failed claiming multi_turn=false. Each context verb now gates on and
reports its own capability. Card schema, capability matrix, the three context
gates, tests, and the lark-agent skill docs are updated in lockstep.
Review follow-up on the agent command tree. A batch of low-risk hardening
fixes; no behavior change for the shipped example provider.
- Terminal-injection: pretty/TSV renderers now sanitize the agent-controlled
State / UpdatedAt / CreatedAt fields (kvValue on pretty rows, stripANSI on
TSV), matching the id/summary/title fields — a malicious provider can no
longer inject CSI/OSC escapes via a forged state or timestamp.
- Nil-safety: `task get --watch` and artifact download return a typed
invalid_response error when a provider hook yields (nil, nil) (a legitimate
Call[*T] result on an empty "data") instead of panicking; an artifact with
neither inline bytes nor a URL no longer writes a 0-byte file.
- Array convention: task/context/agent list normalize a nil slice to [] so an
empty list serializes as [] not null, matching Card.Parameters.
- Error hint: unknown-agent errors keep LookupSpec's scheme-scoped
`agent list <scheme>` hint instead of being flattened to the generic one.
- agent list <scheme> (online path) sets the resolved identity on its
envelope, consistent with the other leaves.
- Comment/doc drift: drop references to the removed Deps probe / Discoverer /
ProviderInfo / resolveProvider symbols; rename Supports(cap) -> capKey.
- Tests: cross-agent isolation in the example store, empty-list [] contract,
State/timestamp sanitization regression, and a real stdout assertion for
context list --jq.
Runtime.CallAPI/CallMultipart now return the response "data" object as
json.RawMessage instead of map[string]any, and provider hooks decode it through
the generic Call[T] / CallUpload[T] helpers: a hook declares the response struct
it expects and the framework unmarshals + classifies errors, instead of poking
at a map. Call[map[string]any] remains available for genuinely dynamic shapes; a
response with no "data" (e.g. a pure write) yields T's zero value and a nil error.
decodeData centralizes the unmarshal + invalid_response classification. cmdRuntime
re-encodes the unwrapped "data" sub-object to raw JSON after CheckResponse. Test
doubles (cmd/agent + example fakeRuntime, card_test fakeRT) updated to the raw
signature; the runtime tests keep the raw-data assertion and add a Call[T]
decode assertion.
The agent scope preflight now mirrors cmd/event's scopeRemediationHint for both
identities, replacing the bespoke replacement-era hint:
- user: the re-auth hint lists ONLY the missing scopes (the open platform
authorizes incrementally, so re-login with just the missing keeps existing
grants — no merge needed). Uses the canonical repo-wide `auth login --scope`
phrasing instead of a one-off Chinese string.
- bot: previously skipped entirely; now checks the app's published TenantScopes
(fetched best-effort via appmeta.FetchCurrentPublished behind a swappable seam;
a fetch failure downgrades to a no-op). A missing scope reports the
developer-console re-publish remediation (the event-style scan-to-enable deep
link lives in cmd/event and is not duplicated here).
preflightScopesForRef keeps its signature (no call-site changes); bot fetch uses
a bounded background context so no ctx param is threaded. Tests cover the
incremental user hint, the bot missing/present/no-scopes branches, and the bot
seam wiring.
`task list` / `context get` carried only {task_id, context_id, state,
is_terminal} — too thin for a caller (especially an AI) to tell which task
to resume without a `task get` per item. Enrich the summary surface so the
list is self-sufficient for triage, aligning with A2A's Task
(status.timestamp + last message).
- TaskSummary: add updated_at + summary (last agent message, or the pending
prompt for input_required; rune-truncated)
- AgentTask: add created_at + updated_at
- ContextSummary: add updated_at + task_count + awaiting_input
- ContextDetail: drop the embedded tasks[]; add updated_at, task_count,
awaiting_input, and active_task (the latest-updated task). Full task
enumeration stays in `agent task list --context-id`.
- task list / context list sort by updated_at desc
- route task list / context list / context get through the content-safety
scan (they now carry untrusted agent text); ANSI-strip + flatten summary
in pretty/TSV
- example provider fills the new fields; tests + lark-agent skill docs updated
Supersede the struct-of-func-fields Provider with a runtime-injection model that
stops leaking framework plumbing to integrators (the Deps{Client, As} an
onboarding author previously had to receive and destructure — the mock didn't
even use it).
Framework (internal/agent):
- Provider is now one declarative value per business domain (scheme): metadata +
a Catalog []AgentSpec (offline-enumerable) XOR an Instance *AgentSpec template,
plus an optional online ListAgents hook. AgentSpec carries per-agent card
metadata, the FileInput/InputRequired flags, and the verb hooks.
- Every hook receives an identity-opaque agent.Runtime (AgentID/IsBot/CallAPI/
CallMultipart) instead of a raw client; the concrete cmdRuntime lives in
cmd/agent (like event's consumeRuntime), so internal/agent no longer depends on
internal/client and the Deps struct is gone. CallMultipart is the centralized,
SafeInputPath-validated file-upload seam that makes file_input deliverable.
- Register takes a Provider (pure-struct validation, fail-fast). LookupSpec
resolves ref→spec fully offline. Capability = wired-hook presence
(DeriveCapabilities), card synthesized by BuildCard (rt=nil ⇒ offline caps +
static metadata; rt!=nil ⇒ best-effort Describe enrichment). Deleted catalog.go,
Deps, Factory, Resolve, NewCard, the zero-Deps probe, and the Discoverer interface.
Command layer (cmd/agent):
- resolveSpec (offline: identity + LookupSpec) then capability nil-gate BEFORE
runtimeFor, so an unsupported verb returns unsupported_capability (exit 2)
before any client is built — uniformly across list/context/artifact (previously
only cancel gated offline). Then runtimeFor + scope preflight + spec.<hook>(rt).
- agent list: catalog enumerates offline (ListCatalog); instance enumerates via
the online ListAgents hook, else reports not-enumerable.
Providers: example is now a declarative Provider() value + plain hooks reading
rt.AgentID() (echo minimal / reporter full differ only by wired fields). Explicit
aggregation in agent/register.go.
Adds cmd/agent runtime tests (CallAPI unwrap/error/transport, IsBot, CallMultipart
SafeInputPath), negative capability-gate tests (send --file / task list / context
get / artifact download), the https-only artifact check, and BuildCard's dynamic
Describe path.
Replace the fat 9-method Provider interface with a struct of function fields,
mirroring the events KeyDefinition / shortcuts Shortcut convention: a provider
wires only the capabilities it supports and leaves the rest nil. This removes the
two things every integrator previously had to keep in sync by hand — the
Capabilities bool matrix and the per-method ErrUnsupported returns.
- Provider is now a struct: core Send/GetTask (mandatory, asserted at Register)
plus optional func fields (ListTasks/CancelTask/context trio/DownloadArtifact/
ListAgents) whose presence == support, plus FileInput/InputRequired flags and
an optional Describe for per-agent card metadata.
- The card capability matrix is DERIVED from which fields are wired
(DeriveCapabilities / BuildCard), so declaration and behavior are single-
sourced and cannot drift. CatalogEntry drops its Capabilities field.
- The command layer gates every optional verb on the nil field and returns a
unified unsupported_capability (exit 2) before any network access; the
ErrUnsupported sentinel and convertUnsupported are deleted. --file is now
capability-gated too (file_input=false ⇒ unsupported before the upload prompt).
- The Discoverer interface becomes the ListAgents field; catalog providers must
wire it (asserted at Register). example expresses echo's minimal set vs
reporter's full set purely by which fields its Factory wires per agent — no
bool matrix, no refusal code. Conformance + tests updated to the new shape.
Mirror the events layering: the framework/SPI stays in internal/agent, the
concrete business providers move to a top-level agent/ package (agent/example/),
and agent/register.go blank-imports each so their init() self-registration runs.
cmd/build.go blank-imports the top-level agent package (alongside events); the
command layer (cmd/agent) no longer wires providers directly.
A test-only blank import keeps the example scheme registered for cmd/agent
tests, which exercise example:echo / example:reporter offline.
Add the `lark-cli agent` command tree: a provider-neutral surface over
remote A2A agents. One constant verb set (list / card / send / task /
context) routes by agent_ref (<scheme>:<agent_id>) to registered
providers; remote agents never grow new top-level commands and their
capabilities are declared in a machine-readable card.
- SPI (internal/agent): Provider interface, registry with fail-fast
registration checks, typed ProviderKind / IdentityType, a closed
Capabilities struct, NewCard single-source card synthesis, and the
9-state task machine aligned with A2A.
- Command surface (cmd/agent): list / card / send / task / context with
default JSON envelopes, meta.next suggested commands, fire + bounded
`--watch --timeout` polling, local all-or-nothing scope preflight,
and capability gating. Two CLI-enforced high-risk-write confirmations:
`send --file` (off-machine upload) needs --yes, and artifact download
refuses to clobber an existing -o target without --force; both return
confirmation_required (exit 10) before any network/write. Artifact
download is SSRF-guarded, https-only and size-capped.
- Typed error contract with stable exit codes and codemeta classification.
- Ignore local-only proof artifacts (tests_e2e/, tests_skill_eval/,
coverage.html).
Both are registered, reachable business domains — `lark-cli attendance
user_tasks query` and `lark-cli mindnotes nodes create` run, and both appear in
`auth login --domain` — but neither had an entry in
service_descriptions.json. GetServiceTitle/GetServiceDescription returned "",
so buildDomainMeta fell back to the typed service spec, which carries one
string for both languages. The result was visible in `--help`: mindnotes
rendered its Chinese description in the English domain list, and attendance
rendered "attendance record query", a lowercase fragment shown in both locales.
Both keep their own scope namespaces (attendance:task:readonly,
mindnote:node:create/read) and stay independent auth domains, so no
auth_domain is set. whiteboard remains the only domain folded into docs.
TestGetDomainMetadata_HasTitleAndDescription could not catch this: it asserts
on buildDomainMeta's output, which the fallback had already made non-empty.
The new reconciliation test walks EmbeddedServicesTyped plus AllShortcuts and
asserts both languages resolve from the config itself, before any fallback.
EmbeddedServicesTyped is the overlay-free parse, so the test does not depend on
what remote_meta.json happens to hold on the machine — a domain that only ever
arrives via remote overlay stays out of its reach.
* refactor: add output emitter contract and differential harness
Introduce a leaf Emitter in internal/output that composes the existing
output primitives (content-safety scan, envelope, jq, format rendering,
notice) behind a single command-scoped port. The emitter is unwired: no
production caller is migrated, so CLI output stays byte-for-byte unchanged.
A differential test harness drives the real legacy entry points
(RuntimeContext.Out/OutRaw/OutFormat/..., WriteSuccessEnvelope and the
pagination formatter) and asserts byte-identical stdout/stderr plus typed
errors, locking behavior before later slices migrate callers.
* refactor: tighten emitter API and cover pagination with real tests
- split Emitter.Success/PartialFailure and drop EmitOptions.OK so a
missing ok flag can no longer silently emit ok:false
- give StreamPage its own StreamOptions (format + pretty) instead of
reusing EmitOptions, making "jq needs aggregation" a compile-time fact
- pin the Emitter jq-error contract (returns error, writes no stderr);
the caller adapter re-emits the legacy stderr line on migration
- add in-package tests driving the real apiPaginate/servicePaginate over
a mock transport: multi-page aggregation, empty-result fallback,
MarkRaw handling, and the business-error raw-response red line
* test: use standard TestFactory harness for pagination tests
Replace the hand-rolled RoundTripper + APIClient construction in the
apiPaginate/servicePaginate tests with cmdutil.TestFactory and its
httpmock.Registry, and isolate LARKSUITE_CLI_CONFIG_DIR to t.TempDir(),
matching the repo's standard HTTP-mocked test convention. Assertions and
coverage (multi-page aggregation, empty-result fallback, MarkRaw, and the
business-error raw-response red line) are unchanged.
* refactor: route success output through the single Emitter port
Migrate the success-output surfaces onto internal/output's Emitter,
byte-for-byte identical (proven by frozen golden diffs and the real
paginate/HandleResponse tests):
- RuntimeContext.Out/OutRaw/OutFormat/OutFormatRaw/OutPartialFailure now
build an Emitter and call Success/PartialFailure; emit and outFormat are
removed. An adapter maps the returned error back to the legacy
outputErrOnce / jq-error stderr / exit-code behavior.
- WriteSuccessEnvelope degrades to a thin Emitter.Success delegate; its 8
callers are unchanged.
- apiPaginate/servicePaginate stream pages via Emitter.StreamPage; the
aggregate and business-error raw-response branches are untouched.
- HandleResponse routes its non-JSON structured-response branch through
Emitter.Success.
Frozen golden fixtures replace the runtime legacy oracles so the
differential harness cannot go self-referential after migration.
* fix: keep _notice on struct payloads in Emitter's unknown-format fallback
printLegacyDataJSON now normalizes via toGeneric first (matching FormatValue), so a struct / named-map payload retains its injected _notice on the unknown-format -> JSON fallback rather than dropping it silently. Add a regression test that fails against the pre-fix path.
* refactor: make the Emitter own write failures and stop mutating inputs
Route every Emitter stdout path through a render-to-buffer-then-copy helper so a marshal/render failure leaves stdout empty and surfaces a typed internal error (with cause), and a stdout write failure is propagated instead of silently swallowed. Leaf writers gain error-returning Write* cores; the legacy Print*/FormatValue wrappers keep their exact behavior for unmigrated callers.
- handleEmitterError now captures every error, not only the jq/safety branches; flip OutRaw's write-error test to assert propagation.
- Clone the map before injecting _notice so a caller's payload is never mutated and an existing _notice is never overwritten.
- Preserve jq's own typed error (validation/api) on a bad expression or runtime failure; only wrap genuine stdout write failures.
- Split tests: normative emitter_contract_test.go vs frozen emitter_legacy_compat_test.go (base SHA recorded, self-update env vars removed).
* fix: satisfy license-header and forbidigo lint on the emitter changes
- Move the base-SHA note below the copyright header in the renamed legacy-compat test so the license-header check sees a valid header at the top.
- Route the leaf wrappers' marshal/format stderr messages through a single legacyStderrf helper (one //nolint:forbidigo) instead of bare os.Stderr, preserving exact legacy behavior for unmigrated direct callers while passing forbidigo; drop the now-unused os imports.
* fix: stop legacy CSV wrappers reporting write failures to stderr
Align FormatAsCSV/FormatAsCSVPaginated and FormatValue/FormatPage's CSV branch with the other leaf wrappers: report only marshal failures, swallow write failures. Previously they emitted a 'csv write error' for the (empty) line and the JSON-fallback write failures that the pre-refactor code ignored, and mislabeled a JSON write failure as a CSV one. Failure-path only; success output is unchanged (golden double-diff still byte-for-byte).
Register approval.instance.status_changed_v4 and approval.task.status_changed_v4 with custom flattened schemas and user-auth pre-consume subscription setup.
Handle approval subscription_type as optional multi-value pre-registration metadata: omitted values register both involved and managed relations, explicit values can be single, comma-separated, or JSON array, and consumers do not unsubscribe on exit.
Report partial approval subscription registration failures with registered and failed relation context while preserving the underlying typed error classification.
Document approval event output fields and subscription semantics, and refresh approval skill references from API metadata.
* fix: unify dry-run output contract
* fix: address dry-run review feedback
* fix(dryrun): tighten preview contract and unify data shape
- transcribe HTTP method verbatim in previews (HEAD/OPTIONS were
reported as GET); reject an empty method in api with a typed error
- unify the dry-run data payload across api/service/shortcut paths:
{api, context?: {app_id, user_open_id}}; drop data.as — the envelope
top-level identity is the single identity source
- mark pretty dry-run stdout with '# dry-run: request not sent' so logs
that drop stderr still show it was a preview
- extract the shared preview builder, collapse PrintDryRunWithFile's
loose params into FileUploadMeta, and fail loudly on nil previews
- revert description-marker identity parsing: stale prose must not
override corrected accessTokens (blocks legal user calls on
images.create); identity gating keys off accessTokens only
- pin the new contracts with tests: verbatim method, three-way context
parity, nil-preview error, empty-context omission, marker line
* docs(agents): add typed-data, faithful-transcription, and contract-test conventions
- typed struct at the boundary over map[string]interface{} threading;
distinct types where values could swap silently (internal/meta.Token)
- transcribe input verbatim in previews/transformations; reject
unhonorable flag combinations with typed errors instead of silently
substituting behavior
- contract tests must fail when the implementation is reverted
* test: migrate dry-run tests grown on main to the envelope format
main gained raw-format dry-run readers while the PR was in flight
(wiki drive export #1802, drive list comments #1845, slash commands,
sheets history, docs fetch, mail draft-send/triage, vc meeting events).
Migrate them to the envelope accessors (clie2e.DryRunGet / data-wrapped
decoders) and drop the now-redundant DryRunData extractions in files
unified on DryRunGet.
---------
Co-authored-by: guokexin.02 <264159873+Tantanz20020918@users.noreply.github.com>
This PR improves the im.message.receive_v1 event output by exposing structural
metadata fields (reply context, sender type, mentions) that were previously only
available in the raw V2 envelope. It also syncs the same structural fields to the legacy
+subscribe --compact pipeline.