Integrating ClickHouse with Linear
The Linear registry item copies independent raw resource readers and ClickHouse destinations into a chkit project.
Install
Section titled “Install”bunx chkit add linearbunx chkit checkbunx chkit generate --name add_linearbunx chkit migrate --applybunx chkit ingest run --tag provider:linearSet LINEAR_API_KEY in the runtime environment before ingestion. Edit queries and raw destinations in src/integrations/linear/sources/ to select fields and configure tables. config.ts supplies source identity and timestamp selection; client.ts handles authenticated requests, and pipeline.ts groups nine independent streams in one workspace pipeline.
createLinearPipeline(config, deps) snapshots configuration and binds injectable HTTP dependencies. The pipeline and all nine table exports remain in index.ts. Destination schemas stay in source modules rather than runtime configuration.
Keep stream IDs and destinations tied to one workspace. The reader does not resolve the token’s workspace identity, so switching workspaces requires new IDs and destinations. Reconcile older observations after adding requested fields.
| Resource | Default ClickHouse table | Records synced | API reference |
|---|---|---|---|
Issues (issues) | linear_issues_raw | Archived and active issue fields, relationship IDs, and fully paginated label names as strings. | POST/graphql |
Comments (comments) | linear_comments_raw | Raw comments selected by their own update times, with nullable issue and other parent references. | POST/graphql |
Projects (projects) | linear_projects_raw | Project metadata and provider relationship references selected by project update times. | POST/graphql |
Project updates (project_updates) | linear_project_updates_raw | Authored progress and health updates with project and author references. | POST/graphql |
Cycles (cycles) | linear_cycles_raw | Cycle dates, descriptions, and team references selected by cycle update times. | POST/graphql |
Users (users) | linear_users_raw | Accessible user metadata, including disabled users, selected by user update times. | POST/graphql |
Teams (teams) | linear_teams_raw | Accessible team metadata selected by team update times. | POST/graphql |
Issue relations (issue_relations) | linear_issue_relations_raw | Complete accessible relation records with type and both directed issue references; no reliable root change filter exists. | POST/graphql |
Issue history (issue_history) | linear_issue_history_raw | Complete accessible history for independently discovered issues; the parent history connection has no timestamp filter. | POST/graphql |
Raw data and resource selection
Section titled “Raw data and resource selection”Issues, comments, projects, project updates, cycles, users, teams, issue relations, and issue history each have their own raw table and checkpoint. Relationships retain provider-shaped ID references. Join resources and calculate metrics later in ClickHouse. Issues retain fully paginated label names as strings; label definitions, workflow states, and project statuses have no separate streams. Bounded inline state/status metadata remains provider fields.
Comments use their own collection, including nullable references for issue and other parent types. History independently discovers all accessible issues before paging their events. Neither reader depends on the issues stream’s execution or checkpoint. Paginated project memberships and other associations are outside the selected resource scope.
bunx chkit ingest run --tag provider:linear --tag resource:commentsbunx chkit ingest run --tag provider:linear --tag resource:issue_historyRepeated tags use AND matching. The default pipeline executes streams sequentially to limit API usage; ordering does not establish a discovery dependency. Expensive history reads can have their own schedule and execution budget.
Sync behavior
Section titled “Sync behavior”Seven root resources use server-side bounded updatedAt windows, Relay cursors, five minutes of overlap, and explicit archive inclusion. Users also include disabled accounts. Comments and project updates use their own update times, independently of their parents. Linear pagination and date filtering document these selection controls.
paginate() owns request retries and continuation cycle detection. Readers validate GraphQL errors and reject partial success. Each timestamp checkpoint advances only after the complete window and destination writes succeed, including an empty window. Failed windows replay from their lower bound. Issues complete all label pages before publishing their label strings.
Issue relations lack a reliable root change filter. History is only exposed through each issue’s history connection, with no timestamp filter. These streams use fullSync(): discover their complete accessible scope on each run and journal success after destination acknowledgement, including empty results. Interrupted full reads restart from the beginning. They ignore date bounds and provide no date-range backfill. See Linear’s API schema.
Deleted records and removed relations remain stored. The API does not promise atomic snapshots, and label renames are not assumed to update issue timestamps. Periodic reconciliation refreshes older timestamp-selected observations:
bunx chkit ingest run --tag provider:linear --tag resource:issues --backfill reconcile-2026-10-05 --from 1970-01-01bunx chkit ingest status --tag provider:linear --jsonUse a new backfill ID for each reconciliation. Timestamp replay reads current observations; the history stream preserves accessible provider events. The raw tables require ClickHouse 25.3 or later. See the installed README for configuration and registry installation.
Upgrading existing data
Section titled “Upgrading existing data”Version 0.2.0 adds eight raw destinations while retaining issue row IDs and the linear.issues stream ID. Existing nested comments in old issue observations remain until those rows are refreshed or migrated deliberately. Initial reads populate the new resource tables; application queries should join those destinations rather than use retained nested comments.
Test the readers
Section titled “Test the readers”bunx chkit add linear --with-testsbun test src/integrations/linear/tests/basic.test.tsChangelog
Section titled “Changelog”Version 0.2.0
- Sync issues, comments, projects, project updates, cycles, users, and teams through independent overlapping updated-time windows, including archived and disabled resources.
- Add independent full-sync streams for issue relations and issue history; discover and paginate history for all accessible issues without relying on parent stream progress or modification times.
- Store each resource in its own raw table with provider IDs and relationship references; retain complete issue labels as strings and leave joins and metrics to ClickHouse.
- Use shared pagination, validated GraphQL responses, request-level retries and sink-acknowledged completion checkpoints; separate configuration, injectable requests, readers, and pipeline wiring.
- Existing issue IDs remain stable; migrate retained nested comment observations deliberately and populate the eight new destinations through their own initial syncs.
Version 0.1.0
- Introduce raw GraphQL issue ingestion.