Integrating ClickHouse with Google Meet
The Google Meet registry item copies six raw resource readers with fixed discovery windows and independently journaled parent and child progress into a chkit project.
Install
Section titled “Install”bunx chkit add google-meet --with-testsbunx chkit checkbunx chkit generate --name add_google_meetbunx chkit migrate --applybunx chkit ingest run --tag provider:google-meetbunx chkit ingest status --tag provider:google-meetSet GOOGLE_MEET_ACCESS_TOKEN with meetings.space.readonly access before ingestion. Review source identity, window settings, schema placement, and limits in src/integrations/google-meet/config.ts. Keep the source identity tied to one Google account; these list responses do not identify the authenticated account. Meet requires user authentication; coverage is limited to records accessible to that user. The reader does not refresh tokens or export an entire workspace automatically. Schema imports do not call Google. See authorization.
| Resource | Default ClickHouse table | Records synced | API reference |
|---|---|---|---|
Conferences (conferences) | google_meet_conferences_raw | Conference records visible to the access token | GET/conferenceRecords |
Transcripts (transcripts) | google_meet_transcripts_raw | Transcript metadata for discovered and retained pending conferences | GET/conferenceRecords, GET/conferenceRecords/{conferenceRecord}/transcripts |
Transcript entries (transcript-entries) | google_meet_transcript_entries_raw | Individually keyed, paged transcript entries for retained conferences | GET/conferenceRecords, GET/conferenceRecords/{conferenceRecord}/transcripts, GET/conferenceRecords/{conferenceRecord}/transcripts/{transcript}/entries |
Participants (participants) | google_meet_participants_raw | Native attendance records, including signed-in, anonymous, and phone participants | GET/conferenceRecords, GET/conferenceRecords/{conferenceRecord}/participants |
Participant sessions (participant-sessions) | google_meet_participant_sessions_raw | Individually keyed join/leave sessions with conference and participant names | GET/conferenceRecords, GET/conferenceRecords/{conferenceRecord}/participants, GET/conferenceRecords/{conferenceRecord}/participants/{participant}/participantSessions |
Recordings (recordings) | google_meet_recordings_raw | Recording metadata and native Drive destination references; no file downloads | GET/conferenceRecords, GET/conferenceRecords/{conferenceRecord}/recordings |
Pipeline and streams
Section titled “Pipeline and streams”One installation exports one pipeline with six independently selected streams: conferences, participants, participant-sessions, recordings, transcripts, and transcript-entries. Each owns its raw table and journal progress. Every child stream discovers its own parent conferences; the session stream also enumerates its own participants. No stream depends on another stream running first or completing successfully. Nested pagination follows parent-scoped API endpoints inside the selected resource stream.
bunx chkit ingest run --tag provider:google-meet --tag resource:conferencesbunx chkit ingest run --tag provider:google-meet --tag resource:participantsbunx chkit ingest run --tag provider:google-meet --tag resource:participant-sessionsbunx chkit ingest run --tag provider:google-meet --tag resource:recordingsbunx chkit ingest run --tag provider:google-meet --tag resource:transcriptsbunx chkit ingest run --tag provider:google-meet --tag resource:transcript-entriespipeline.ts wires destinations and resource streams; sources/resources.ts shares the provider state machine, and client.ts handles requests. sourceId scopes raw rows and query state; streamPrefix determines journal IDs. Defaults preserve google-meet.* checkpoints. Another installation needs distinct values for both. The pipeline factory binds configuration and injectable request dependencies while using the exported destinations. Edit database in config.ts before schema imports to change table placement.
The default pipeline runs one stream and request at a time. Filtered runs give resources separate schedules and execution budgets.
Sync behavior
Section titled “Sync behavior”Each stream independently discovers completed conferences through fixed end_time windows and ongoing conferences through end_time IS NULL. End-time discovery includes long-running calls that began before the lookback. The initial range defaults to 30 days, later discovery overlaps seven days, and each cycle advances by at most windowDays.
All child streams retain parent conferences until provider expiry and revisit them on completed cycles. Empty child lists, generated files, late joins, and subsequently populated leave times stay eligible after discovery advances, including outside the discovery overlap. Attendance readers fetch complete participant/session lists rather than filtering on attendance times. This is polling; Meet provides no modification/change token. Recordings/transcripts/entries have no date filters, and participant/session time filters do not select updates. See conference filters and artifact behavior.
Checkpoints and recovery
Section titled “Checkpoints and recovery”The journal stores fixed cycle bounds and resource-specific parent and child-page positions. Session readers retain one participant-name page and nested session progress; entry readers retain one transcript page and nested entry progress. Candidate progress commits after its covering rows load; failed loads retain the preceding position. Rerun after interruption or budget exhaustion to continue. Rejected page tokens replay their scoped collection once, and repeated or malformed tokens fail visibly.
Saved recovery counts cover each unfinished discovery phase and child collection, including nested sessions and entries. They clear at the acknowledged terminal boundary, so pauses and nonterminal replay pages do not renew the allowance. A second rejection fails across executions. Adjust maxChunks in config.ts, execution duration, or polling frequency, then review coverage before explicitly migrating saved state or using a new stream identity for a fresh scan.
Streams permit 200 chunks per run and retain at most 1,000 conferences. Edit maxChunks and maxPendingConferences in config.ts; an oversized pending queue fails instead of losing parents. Source/window changes require explicit checkpoint migration or a new stream identity. Native JSON requires ClickHouse 25.3 or later. Schedule repeat runs externally with one ingestion process per destination at a time.
Google removes conference records and API transcript entries 30 days after a call ends. Expiry during unfinished work fails with a coverage-gap message. Previously checked parents leave the queue on expiry, while stored raw rows remain. Expired or inaccessible resources do not become inferred deletions. See conference retention and artifact retention.
Data and historical ranges
Section titled “Data and historical ranges”Raw identities are [sourceId, provider.name]. Provider objects keep their native fields. Child rows add conference_name; session rows also add participant_name, and entry rows add transcript_name. Join raw resources through their full provider names in ClickHouse.
Participants are attendance records for one call, with native signed-in, anonymous, or phone identities. Sessions represent individual joins/leaves from a device. Recordings contain metadata, state, and Drive references without downloading files. Transcripts contain metadata and document references; individual entry rows contain speech text and participant references. Attendance totals, assembled transcripts, speaker enrichment, and other joined views belong in ClickHouse. See participants, sessions, and recordings.
Entries stream page by page into google_meet_transcript_entries_raw. The API entries do not capture later edits to the separate Google Docs transcript file.
Version 0.2.0 adds participant, session, recording, and entry destinations and stores new transcript metadata without embedding every entry in one row. Existing tables and conference/transcript stream IDs remain available. Added streams start their own discovery/checkpoints. Row identities now include the source label; keep older datasets as archives or migrate their identities deliberately. See the installed README for upgrade details.
Explicit backfill bounds constrain conference end times and skip ongoing-call discovery. Reuse the same backfill ID and bounds to resume; expired API history remains unavailable. Bounds select conference end times, not child creation, attendance, or update times. Child lists remain parent-scoped and complete for selected and retained conferences.
bunx chkit ingest run --tag provider:google-meet --backfill october --from 2026-10-01T00:00:00Z --to 2026-10-05T00:00:00ZFixture verification
Section titled “Fixture verification”bun test src/integrations/google-meet/tests/basic.test.tsFixtures cover all six resources independently and together, nested session/entry resumption, late attendance and recordings from retained parents, source failure, lost sink acknowledgement, changed replay content, persisted token-recovery debt, isolated backfill bounds, and expiry gaps.
Changelog
Section titled “Changelog”Version 0.2.0
- Sync six independent raw resources in one installation pipeline: conferences, participants, participant sessions, recordings, transcript metadata, and individual transcript entries.
- Checkpoint fixed conference discovery windows and retain pending conferences until expiry to revisit late artifacts and attendance changes independently of parent updates.
- Resume acknowledged discovery and nested child pages with bounded token recovery, sink acknowledgement, and visible expiry gaps.
- Keep recording metadata and Drive references without downloading files; keep child rows separate with provider-name join keys and source-scoped row IDs. Migrate legacy transcript identities and embedded entries deliberately when upgrading.
- Separate editable installation configuration, injectable HTTP clients, and resource readers; preserve existing conference/transcript/entry stream identities and checkpoint scope.
Version 0.1.0
- Introduce raw conference and transcript ingestion with transcript entries nested in each transcript observation.