Files
Colby McHenry 1a678aabc7 docs(telemetry): tell the truth about where events are stored (CG-15)
The telemetry docs are a privacy contract, and they still described a
managed analytics store that no longer receives anything. Replace that
with what actually happens now — events land in our own D1 database on
Cloudflare, the endpoint makes no outbound requests, raw events are
purged after 90 days and only anonymous daily rollups outlive them.
This strengthens the guarantee rather than restating it: there is no
second party to share with.

- TELEMETRY.md: new "Where it is stored" section; the never-collected
  IP bullet no longer leans on a vendor-side setting to hold.
- docs/design/telemetry.md: ingest section rewritten around D1 + the
  nightly rollup/retention cron; volume math redone on Workers Paid and
  the D1 quota (storage, not writes, is what sets the 90-day window);
  new section documenting the dashboard worker and cross-linking it.
- Fixed three drifts from the worker allowlist the sweep surfaced:
  schema_version was still 1, client_name/client_version was still
  marked "plumbing to add" though session.ts passes it today, and the
  legacy sqlite_backend field the worker still accepts was undocumented.
- telemetry-worker/README.md: step 6 claimed a repo-wide grep came back
  clean, which this runbook itself falsifies. Added step 7 — deleting
  the runbook is what makes that grep true, and is the completion check.
- smoke-cutover.sh: the vendor guarantee is now asserted by class
  (no analytics-ingest endpoint referenced) rather than by one vendor's
  name, so it keeps working once the name is gone. Verified it still
  catches a planted forwarding URL. 61/61 pass.

Retention is documented as 90 days, not the 180 in the task notes: 180
days of raw events exceeds D1's 10 GB per-database cap, and the code
purges at 90.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-28 22:10:43 -05:00

5.8 KiB

Telemetry

CodeGraph collects a small set of anonymous usage statistics — which commands and tools get used, which languages get indexed, which agents drive usage — so we can tell which of the 20+ languages and 8 agent integrations deserve the most work. This page is the complete list of what is collected. If a field isn't on this page, it isn't collected; the ingest endpoint enforces this list as an allowlist and is itself public, auditable code in this repository.

Turning it off

Any of these works, permanently:

codegraph telemetry off        # stores your choice (and deletes any unsent data)
export CODEGRAPH_TELEMETRY=0   # per-shell / per-CI override
export DO_NOT_TRACK=1          # the cross-tool standard — always honored

codegraph telemetry status shows the current state, what decided it, and your machine ID. The interactive installer (codegraph install) asks up front with a visible default-on toggle and never re-asks. If you never saw the installer (e.g. npx straight into init), a one-line notice is printed to stderr before the first time anything is sent.

Off means off: when disabled, CodeGraph records nothing, opens no connection to the telemetry endpoint, and sends no "opted out" ping.

Separately from telemetry, the MCP server checks GitHub for a newer release in the background (at most once a day) so it can tell you an update exists — it fetches a version number and sends nothing about you or your machine. DO_NOT_TRACK=1 disables this check too; to turn off only the update check, use CODEGRAPH_NO_UPDATE_CHECK=1.

What is collected

Every payload carries this envelope:

field example notes
machine_id b3a8c1… random UUID minted on first send — derived from nothing
codegraph_version 0.9.9
os / arch darwin / arm64 platform identifiers only
node_major 22 major version only
ci false whether the CI env var was set
schema_version 2 bumped when this page changes (v2 dropped the index event's sqlite_backend field)

And one of four events:

  • install — when codegraph install configures agents: which agents (["claude","cursor",…]), global vs project-local, and whether it was a fresh install, an upgrade, or a re-run.
  • index — when a full index completes: the language names present (e.g. ["typescript","go"]), the file count as a coarse bucket (<100, 100-1k, 1k-10k, 10k+), and the duration as a bucket (<10s, 10-60s, 1-5m, 5m+).
  • usage_rollup — one line per day per tool: the tool or CLI command name (e.g. codegraph_explore, init), how many times it ran, how many errored, and — for MCP tools — the connecting agent's name and version from the MCP handshake (e.g. Claude Code 2.1). The Claude Code prompt hook also counts its gate decision (fired fully, fired as a hint, or did nothing — fixed counter names like prompt-hook-gate-medium-segment); the prompt itself is never read, stored, or sent.
  • uninstall — when codegraph uninstall/uninit runs: which agents were removed.

Usage is aggregated locally into daily totals before anything is sent — there is no per-call event stream, and nothing is sent in real time.

What is never collected

  • No source code. No file paths, file names, directory names, repository names or URLs, symbol names, search queries, or anything else derived from the contents of an indexed project.
  • No IP addresses. The ingest endpoint never reads, logs, or stores the client IP — and there is no analytics vendor downstream that could. No geolocation.
  • No fingerprinting. The machine ID is a random UUID stored in ~/.codegraph/telemetry.json — delete that file (or run codegraph telemetry off, then on) and the old ID is gone forever, with no way to reconnect it.
  • No personal data. No usernames, hostnames, emails, or environment variables.

How it travels

Events POST to telemetry.getcodegraph.com — a first-party endpoint whose complete source lives in telemetry-worker/ in this repository. It validates every event and property against the allowlist above (anything else is dropped), never reads the client IP, and rate-limits per machine ID. Sends are fire-and-forget with a short timeout: offline or air-gapped machines buffer a bounded local file (256 KB cap) and never retry-loop, log errors, or slow a command down. Telemetry never adds latency to MCP tool calls — recording is an in-memory counter.

Where it is stored

Accepted events are written to our own database on Cloudflare (D1) and go nowhere else. No third-party analytics vendor receives any of this data, because the ingest endpoint makes no outbound requests at all — its source is the entire path your events take, and there is nothing after it. This is a stronger guarantee than a promise not to share: there is no second party to share with.

What is kept is checkable rather than asserted. The storage schema — telemetry-worker/migrations/0001_init.sql, checked in beside the endpoint that writes it — is the complete list of what a row can hold, with a comment on every column.

Individual events are deleted after 90 days. What outlives them is anonymous daily totals: counts per day of things like operating system, version, and language, plus which days each machine ID was active so returning-user numbers survive. No event details, and still nothing that identifies a person or a codebase.

The engineering contract behind all of this — including the rule that schema changes must update this page, the client, and the public endpoint in one PR — is in docs/design/telemetry.md.