fix(webapp,clickhouse): keep the rest of a ClickHouse batch when one run or span has un-ingestable JSON#4358
fix(webapp,clickhouse): keep the rest of a ClickHouse batch when one run or span has un-ingestable JSON#4358ericallam wants to merge 7 commits into
Conversation
|
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughClickHouse event and run replication now recover from unparseable JSON by isolating invalid rows and preserving parseable rows in the batch. Both paths accept per-insert settings overrides, expose recovery metrics, and track landed or skipped rows. New integration tests cover mixed event and task-run batches. Benchmark tooling adds configurable poisoned payloads, producer statistics, event-loop utilization measurements, and healthy-versus-poisoned scenarios. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
🧹 Nitpick comments (2)
apps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts (1)
161-165: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winCounter name
rows_isolatedcounts batches, not rows. The metric is incremented once per recovered batch (this._rowsIsolatedCounter.add(1, …),unit: "batches", description says "Batches that recovered…"), yet the name reads as a row count and sits next to the genuine row metric_permanentlyDroppedRows. Once exported to Prometheus this becomesingest_flush_rows_isolated_total, which an operator will very likely read as isolated rows. Renaming now (e.g.ingest.flush.batches_row_isolated) is cheap; after dashboards/alerts are built it is not.As per coding guidelines: be aware that the OTLP→Prometheus exporter adds suffixes, so account for how the metric name reads in queries/dashboards.
Source: Coding guidelines
apps/webapp/app/services/runsReplicationService.server.ts (1)
1216-1221: 🩺 Stability & Availability | 🔵 TrivialRuns replication exposes isolation/drop counts only as getters — no OTEL metric. The sibling
ClickhouseEventRepositoryemitsingest.flush.rows_isolated(andingest.flush.batches_dropped) viaCounters, but here_rowIsolationRecoveriesand_permanentlyDroppedBatchesare just fields read by tests. Since this is precisely the "un-ingestable JSON silently skipped / batch dropped" path, consider wiring both into the existing_meter(as counters with a boundedtable/contextLabelattribute) so production drops and isolations are alertable, not just test-observable.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro Plus
Run ID: 8202e663-cbe4-454e-a1c6-c054e86de979
📒 Files selected for processing (7)
.server-changes/fix-runs-list-batch-drop-on-unparseable-json.mdapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.tsapps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationService.part10.test.ts
📜 Review details
⏰ Context from checks skipped due to timeout. (15)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (2, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (10, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (6, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (7, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (11, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (9, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (5, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (4, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (12, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (8, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (3, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (1, 12)
- GitHub Check: e2e-webapp / 🧪 E2E Tests: Webapp
- GitHub Check: typecheck / typecheck
- GitHub Check: Detect changes
🧰 Additional context used
📓 Path-based instructions (12)
**/*.{ts,tsx}
📄 CodeRabbit inference engine (.github/copilot-instructions.md)
**/*.{ts,tsx}: Use types over interfaces for TypeScript
Avoid using enums; prefer string unions or const objects instead
**/*.{ts,tsx}: Prefer static imports over dynamicimport(); use dynamic imports only for unresolvable circular dependencies, genuine performance code splitting, or conditional runtime loading.
Import Trigger.dev tasks from@trigger.dev/sdk; never use@trigger.dev/sdk/v3or deprecatedclient.defineJob.
Add agentcrumbs while writing code using approved namespaces; mark lines with//@Crumbsor blocks with `// `#region` `@crumbs, and strip them before merging.
Files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
{packages/core,apps/webapp}/**/*.{ts,tsx}
📄 CodeRabbit inference engine (.github/copilot-instructions.md)
Use zod for validation in packages/core and apps/webapp
Files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
**/*.{ts,tsx,js,jsx}
📄 CodeRabbit inference engine (.github/copilot-instructions.md)
Use function declarations instead of default exports
Files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
**/*.{test,spec}.{ts,tsx}
📄 CodeRabbit inference engine (.github/copilot-instructions.md)
Use vitest for all tests in the Trigger.dev repository
**/*.{test,spec}.{ts,tsx}: Use Vitest exclusively and never mock dependencies; use Testcontainers for integration dependencies.
Place test files next to the source files they test.
Files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.ts
**/*.ts
📄 CodeRabbit inference engine (.cursor/rules/otel-metrics.mdc)
**/*.ts: When creating or editing OTEL metrics (counters, histograms, gauges), ensure metric attributes have low cardinality by using only enums, booleans, bounded error codes, or bounded shard IDs
Do not use high-cardinality attributes in OTEL metrics such as UUIDs/IDs (envId, userId, runId, projectId, organizationId), unbounded integers (itemCount, batchSize, retryCount), timestamps (createdAt, startTime), or free-form strings (errorMessage, taskName, queueName)
When exporting OTEL metrics via OTLP to Prometheus, be aware that the exporter automatically adds unit suffixes to metric names (e.g., 'my_duration_ms' becomes 'my_duration_ms_milliseconds', 'my_counter' becomes 'my_counter_total'). Account for these transformations when writing Grafana dashboards or Prometheus queries
Files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
apps/webapp/**/*.{ts,tsx}
📄 CodeRabbit inference engine (.cursor/rules/webapp.mdc)
apps/webapp/**/*.{ts,tsx}: Access environment variables through theenvexport ofenv.server.tsinstead of directly accessingprocess.env
Use subpath exports from@trigger.dev/corepackage instead of importing from the root@trigger.dev/corepathDo not reintroduce the removed v1 execution path;
RunEngineVersion.V1branches may only reject or finalize gracefully so v3 clients receive a clean 4xx, never a 5xx.
Files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
apps/webapp/**/*.test.{ts,tsx}
📄 CodeRabbit inference engine (.cursor/rules/webapp.mdc)
Do not import
env.server.tsdirectly or indirectly into test files; instead pass environment-dependent values through options/parameters to make code testable
Files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.ts
apps/**/*.{ts,tsx}
📄 CodeRabbit inference engine (AGENTS.md)
For apps, use
typecheckfor verification and never usebuildas the correctness check.
Files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
apps/webapp/**/*.{test,spec}.{ts,tsx}
📄 CodeRabbit inference engine (apps/webapp/CLAUDE.md)
Test files must not import
app/env.server.ts; pass configuration as options instead.
Files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.ts
apps/webapp/app/**/*.{ts,tsx}
📄 CodeRabbit inference engine (apps/webapp/CLAUDE.md)
apps/webapp/app/**/*.{ts,tsx}: For dashboard changes, visually verify the running Remix app with Chrome DevTools MCP, using snapshots, screenshots, interaction, and console-message checks as appropriate.
UseuseCallbackanduseMemoonly for context provider values, expensive derived data used as a dependency, or stable references required by dependency arrays; do not wrap ordinary event handlers or trivial computations.
Use named constants for sentinel or placeholder values instead of scattering raw string literals across comparisons.
Files:
apps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
apps/webapp/app/**/*.ts
📄 CodeRabbit inference engine (apps/webapp/CLAUDE.md)
apps/webapp/app/**/*.ts: Never userequest.signalto detect client disconnects. UsegetRequestAbortSignal()fromapp/services/httpAsyncStorage.server.ts, which is wired to Express response close events.
Access environment variables through theenvexport fromapp/env.server.ts; never useprocess.envdirectly.
Always use PrismafindFirstinstead offindUnique.
Always use the$transactionhelper from~/db.server, never callprisma.$transactionor$replica.$transactiondirectly. Pass isolation levels as strings, useSerializablefor correctness-critical read-then-write invariants, and guard possibly undefined helper results when a definite value is required.
Files:
apps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
apps/webapp/app/v3/**/*.ts
📄 CodeRabbit inference engine (apps/webapp/CLAUDE.md)
New code must target Run Engine V2 through the singleton in
app/v3/runEngine.server.ts; do not reintroduce V1 execution paths. V1 branches may only reject or finalize gracefully with a clean 4xx.
Files:
apps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
🧠 Learnings (24)
📚 Learning: 2026-05-14T14:54:39.095Z
Learnt from: ericallam
Repo: triggerdotdev/trigger.dev PR: 3545
File: .server-changes/agent-view-sessions.md:10-10
Timestamp: 2026-05-14T14:54:39.095Z
Learning: In the `trigger.dev` repository, do not flag inconsistent dot vs slash notation in route/path strings inside `.server-changes/*.md` files. These markdown files are consumed verbatim into the changelog, so the mixed notation (e.g., `resources.orgs.../runs.$runParam/...`) is intentional and should be preserved as-is.
Applied to files:
.server-changes/fix-runs-list-batch-drop-on-unparseable-json.md
📚 Learning: 2026-03-22T13:26:12.060Z
Learnt from: ericallam
Repo: triggerdotdev/trigger.dev PR: 3244
File: apps/webapp/app/components/code/TextEditor.tsx:81-86
Timestamp: 2026-03-22T13:26:12.060Z
Learning: In the triggerdotdev/trigger.dev codebase, do not flag `navigator.clipboard.writeText(...)` calls for `missing-await`/`unhandled-promise` issues. These clipboard writes are intentionally invoked without `await` and without `catch` handlers across the project; keep that behavior consistent when reviewing TypeScript/TSX files (e.g., usages like in `apps/webapp/app/components/code/TextEditor.tsx`).
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
📚 Learning: 2026-03-22T19:24:14.403Z
Learnt from: matt-aitken
Repo: triggerdotdev/trigger.dev PR: 3187
File: apps/webapp/app/v3/services/alerts/deliverErrorGroupAlert.server.ts:200-204
Timestamp: 2026-03-22T19:24:14.403Z
Learning: In the triggerdotdev/trigger.dev codebase, webhook URLs are not expected to contain embedded credentials/secrets (e.g., fields like `ProjectAlertWebhookProperties` should only hold credential-free webhook endpoints). During code review, if you see logging or inclusion of raw webhook URLs in error messages, do not automatically treat it as a credential-leak/secrets-in-logs issue by default—first verify the URL does not contain embedded credentials (for example, no username/password in the URL, no obvious secret/token query params or fragments). If the URL is credential-free per this project’s conventions, allow the logging.
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
📚 Learning: 2026-05-18T08:21:27.694Z
Learnt from: d-cs
Repo: triggerdotdev/trigger.dev PR: 3632
File: apps/webapp/sentry.server.ts:4-21
Timestamp: 2026-05-18T08:21:27.694Z
Learning: When handling Prisma error P1001 ("Can't reach database server") in TypeScript, don’t assume a single error shape. Prisma can surface P1001 via two different error classes/fields: `PrismaClientKnownRequestError` exposes it as `err.code === "P1001"` (common during mid-query connection drops), while `PrismaClientInitializationError` exposes it as `err.errorCode === "P1001"` (common on client startup failure). Therefore, predicates should use `err.code === "P1001" || err.errorCode === "P1001"`. Do not flag `err.code === "P1001"` as “unreachable/never matches,” as it is expected in production.
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
📚 Learning: 2026-05-18T08:21:27.694Z
Learnt from: d-cs
Repo: triggerdotdev/trigger.dev PR: 3632
File: apps/webapp/sentry.server.ts:4-21
Timestamp: 2026-05-18T08:21:27.694Z
Learning: When handling Prisma errors for P1001 ("Can't reach database server"), do not assume it only appears under a single property name. Prisma may surface P1001 via either `PrismaClientKnownRequestError` (`err.code === "P1001"`, e.g., mid-query connection drops) or `PrismaClientInitializationError` (`err.errorCode === "P1001"`, e.g., client startup connection failure). To reliably detect the condition, check `err.code === "P1001" || err.errorCode === "P1001"`, and avoid review rules that would incorrectly flag `err.code === "P1001"` as unreachable/never-matching.
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
📚 Learning: 2026-06-13T19:53:13.759Z
Learnt from: ericallam
Repo: triggerdotdev/trigger.dev PR: 3937
File: packages/trigger-sdk/skills/realtime-and-frontend/SKILL.md:258-260
Timestamp: 2026-06-13T19:53:13.759Z
Learning: When reviewing code that uses `trigger.dev/react-hooks`’s `useRealtimeRun`, preserve the call signature where the first argument is the full realtime handle object (not `handle.id`). This is intentional to maintain type-safety and is consistent with the official docs; do not suggest changing the first argument from the handle object to `handle.id`.
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
📚 Learning: 2026-06-17T17:13:49.929Z
Learnt from: matt-aitken
Repo: triggerdotdev/trigger.dev PR: 3948
File: apps/webapp/app/routes/_app.orgs.$organizationSlug.projects.$projectParam.env.$envParam.bulk-actions.$bulkActionParam/route.tsx:48-62
Timestamp: 2026-06-17T17:13:49.929Z
Learning: In triggerdotdev/trigger.dev, within `dashboardLoader`/`dashboardAction` (or similar context resolver code) whenever you resolve an organization ID from an organization slug for RBAC/enterprise authorization scope, always read from the primary Prisma client (`prisma`), not `$replica`. Using `$replica` can hit replica-lag and cause the RBAC lookup/authorization to run without the correct org scope (bypassing intended role enforcement). Implement the slug→org lookup with `prisma.organization.findFirst(...)` (or equivalent primary-client query) and add an inline comment documenting why the primary client is required (replica lag could lead to unscoped RBAC checks).
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
📚 Learning: 2026-06-23T13:04:21.413Z
Learnt from: carderne
Repo: triggerdotdev/trigger.dev PR: 4023
File: apps/webapp/app/services/upsertBranch.server.ts:14-18
Timestamp: 2026-06-23T13:04:21.413Z
Learning: In TypeScript, it’s valid to `import { type X }` and then use `typeof X` in a type-only position, e.g. `type Alias = z.infer<typeof X>`. The `type` modifier suppresses the runtime import, but the type checker still has the full exported type so `z.infer<typeof X>` can resolve correctly. In code reviews, don’t flag this as a TypeScript compile error as long as `typeof X` is used in a type context (e.g., with `z.infer`, `type` aliases, generics), not as a runtime value.
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
📚 Learning: 2026-05-07T12:25:18.271Z
Learnt from: d-cs
Repo: triggerdotdev/trigger.dev PR: 3531
File: apps/webapp/test/sentryTraceContext.server.test.ts:9-47
Timestamp: 2026-05-07T12:25:18.271Z
Learning: In the triggerdotdev/trigger.dev webapp test suite, it is acceptable to leave `createInMemoryTracing()` calls that register a global `NodeTracerProvider` without `afterEach`/`afterAll` teardown. Do not flag this as a test-ordering risk when the code follows the established pattern used across webapp tests (e.g., replication service/benchmark/backfiller tests). This is considered safe because `trace.getActiveSpan()` when called outside a `context.with(...)` block reads `AsyncLocalStorage.getStore()` (undefined when no `run()` scope exists), so it falls back to `ROOT_CONTEXT` with no attached span—regardless of which provider is registered.
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.ts
📚 Learning: 2026-05-28T20:02:10.647Z
Learnt from: myftija
Repo: triggerdotdev/trigger.dev PR: 3772
File: apps/webapp/test/findOrCreateBackgroundWorker.test.ts:1-1
Timestamp: 2026-05-28T20:02:10.647Z
Learning: In the triggerdotdev/trigger.dev monorepo, for the `apps/webapp` package use the established convention of storing Vitest tests (unit, integration, and e2e) under `apps/webapp/test/` rather than colocating them next to source files. Do not flag files located in `apps/webapp/test/` as violating any rule that says to colocate tests with source.
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.ts
📚 Learning: 2026-05-12T21:04:05.815Z
Learnt from: ericallam
Repo: triggerdotdev/trigger.dev PR: 3542
File: apps/webapp/app/components/sessions/v1/SessionStatus.tsx:1-3
Timestamp: 2026-05-12T21:04:05.815Z
Learning: In this Remix + TypeScript codebase, do not flag a server/client boundary violation when a file imports only types from a module matching `*.server`.
Specifically, it’s safe to import types using `import type { Foo } from "*.server"` or `import { type Foo } from "*.server"` because TypeScript erases type-only imports at compile time and they emit no JavaScript, so they won’t cross the Remix server/client bundle boundary.
Only raise the boundary concern for value imports (e.g., `import { Foo }` without `type`, or `import Foo`), since those produce JavaScript output.
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
📚 Learning: 2026-06-25T18:21:51.905Z
Learnt from: carderne
Repo: triggerdotdev/trigger.dev PR: 4039
File: apps/webapp/app/routes/invite-revoke.tsx:0-0
Timestamp: 2026-06-25T18:21:51.905Z
Learning: During the Zod v4 migration in the triggerdotdev/trigger.dev webapp, ensure any imports from `conform-to/zod` use the Zod-4 subpath: `conform-to/zod/v4` (e.g., `import { parseWithZod } from "conform-to/zod/v4"`). Do not import from the package root `conform-to/zod`, because it is the Zod 3 implementation and may load Zod-3-only symbols (e.g., `ZodBranded`, `ZodEffects`), which can throw at module load (notably with `zod4.4.3`). This should be enforced across `apps/webapp/**/*` where helpers like `parseWithZod` and `conformZodMessage` are used.
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
📚 Learning: 2026-07-03T17:10:21.498Z
Learnt from: 0ski
Repo: triggerdotdev/trigger.dev PR: 4148
File: apps/webapp/app/models/orgMember.server.ts:149-168
Timestamp: 2026-07-03T17:10:21.498Z
Learning: In triggerdotdev/trigger.dev, `User.email` (Prisma schema: `internal-packages/database/prisma/schema.prisma`) currently does NOT use `citext` and does NOT have a `lower(email)` functional unique index. Therefore, do not introduce Prisma queries like `where: { email: { equals: <value>, mode: "insensitive" } }` (or any case-insensitive lookup) against `User.email`, because it can force sequential scans of the `users` table under load. During review, ensure email is normalized (e.g., lowercased/trimmed) before both writes and subsequent lookups, and if true case-insensitive behavior/uniqueness is required, implement it via a separate app-wide migration (e.g., switch to `citext` and/or add a functional unique index with backfill) rather than bolting it onto individual feature PRs.
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
📚 Learning: 2026-05-18T14:40:02.173Z
Learnt from: ericallam
Repo: triggerdotdev/trigger.dev PR: 3658
File: packages/core/src/v3/realtimeStreams/manager.test.ts:1-147
Timestamp: 2026-05-18T14:40:02.173Z
Learning: In the triggerdotdev/trigger.dev repo, the policy “Never mock anything — use testcontainers instead” should only be enforced for integration tests that interact with real external services (e.g., Redis, Postgres) via actual infrastructure. For unit tests that exercise pure in-memory logic (e.g., cache semantics) it is OK to stub collaborators such as `ApiClient` using Vitest (`vi.fn()`) to assert call counts or control behavior. Do not flag `vi.fn()`-based `ApiClient` stubs in unit tests as violations of the testcontainers policy.
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.ts
📚 Learning: 2026-06-04T18:16:35.386Z
Learnt from: nicktrn
Repo: triggerdotdev/trigger.dev PR: 3836
File: apps/supervisor/src/backpressure/backpressureMonitor.ts:3-5
Timestamp: 2026-06-04T18:16:35.386Z
Learning: When reviewing TypeScript in this repo, apply the rule “prefer type aliases over interfaces” only to data/object shapes and union/intersection type modeling. If an interface is being used as a behavioral contract for collaborators to implement (e.g., method-shape interfaces that define required behavior, such as `BackpressureLogger` / `BackpressureSignalSource` in `apps/supervisor/src/backpressure/backpressureMonitor.ts`), keep it as an `interface` and do not flag it as a type-alias-vs-interface violation.
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
📚 Learning: 2026-06-09T17:58:04.699Z
Learnt from: 0ski
Repo: triggerdotdev/trigger.dev PR: 3879
File: apps/webapp/app/models/vercelIntegration.server.ts:619-630
Timestamp: 2026-06-09T17:58:04.699Z
Learning: In this codebase, outbound raw `fetch` calls should typically rely on Node/undici’s default request timeout (about ~300s) rather than adding a per-call `AbortController` + `setTimeout` wrapper inside individual functions (e.g. in files like `apps/webapp/app/models/vercelIntegration.server.ts`). During code review, do not flag the absence of a per-call timeout on a single `fetch` as an issue; if per-call timeouts are needed, they should be implemented via a codebase-wide convention (e.g., a shared fetch wrapper or documented pattern) rather than ad-hoc per-function changes.
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.tsapps/webapp/test/runsReplicationBenchmark.producer.tsapps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
📚 Learning: 2026-06-16T09:19:47.637Z
Learnt from: d-cs
Repo: triggerdotdev/trigger.dev PR: 3960
File: apps/webapp/test/prismaInfrastructureErrorCapture.test.ts:0-0
Timestamp: 2026-06-16T09:19:47.637Z
Learning: In this repo’s Vitest setup, `vitest.config.ts` uses `globals: true`, so identifiers like `vi`, `describe`, `it`, and `expect` are available as globals in Vitest test files. During code review, do not flag missing `vi`/`describe`/`it`/`expect` imports as a runtime error or correctness issue when they’re used in `*.test.ts/tsx` or `*.spec.ts/tsx` files. Explicit imports are still preferred for consistency, but they’re not required for runtime behavior.
Applied to files:
apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.tsapps/webapp/test/runsReplicationService.part10.test.tsapps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.ts
📚 Learning: 2026-03-26T09:02:07.973Z
Learnt from: myftija
Repo: triggerdotdev/trigger.dev PR: 3274
File: apps/webapp/app/services/runsReplicationService.server.ts:922-924
Timestamp: 2026-03-26T09:02:07.973Z
Learning: When parsing Trigger.dev task run annotations in server-side services, keep `TaskRun.annotations` strictly conforming to the `RunAnnotations` schema from `trigger.dev/core/v3`. If the code already uses `RunAnnotations.safeParse` (e.g., in a `#parseAnnotations` helper), treat that as intentional/necessary for atomic, schema-accurate annotation handling. Do not recommend relaxing the annotation payload schema or using a permissive “passthrough” parse path, since the annotations are expected to be written atomically in one operation and should not contain partial/legacy payloads that would require a looser parser.
Applied to files:
apps/webapp/app/services/runsReplicationService.server.ts
📚 Learning: 2026-04-20T14:50:16.440Z
Learnt from: ericallam
Repo: triggerdotdev/trigger.dev PR: 3417
File: apps/webapp/app/services/sessionsReplicationService.server.ts:224-231
Timestamp: 2026-04-20T14:50:16.440Z
Learning: In Trigger.dev’s replication services (e.g., sessionsReplicationService.server.ts and runsReplicationService.server.ts), the “acknowledge-before-flush” behavior is intentional. The `_latestCommitEndLsn` should be updated at Postgres commit time and acknowledged on a periodic interval (via methods like `#acknowledgeLatestTransaction`) without waiting for ClickHouse batch flush to complete. Reviewers should not flag this as a durability/ordering bug; it is an established project-wide at-least-once delivery trade-off used across both runs and sessions replication services.
Applied to files:
apps/webapp/app/services/runsReplicationService.server.ts
📚 Learning: 2026-04-20T15:08:49.959Z
Learnt from: ericallam
Repo: triggerdotdev/trigger.dev PR: 3417
File: apps/webapp/app/services/sessionsReplicationService.server.ts:204-215
Timestamp: 2026-04-20T15:08:49.959Z
Learning: For replication services in `apps/webapp/app/services/*ReplicationService.server.ts`, keep the `ConcurrentFlushScheduler` deduplication key shape consistent across the related services (e.g., sessions vs runs) by using the same `${item.event}_${item.session.id}` / `${item.event}_${item.run.id}` pattern. If the key format ever needs to change (such as keying only by session/run id), make the update in all related replication services together—never in just one—so deduplication behavior stays aligned across services.
Applied to files:
apps/webapp/app/services/runsReplicationService.server.ts
📚 Learning: 2026-05-05T09:38:02.512Z
Learnt from: d-cs
Repo: triggerdotdev/trigger.dev PR: 3523
File: apps/webapp/app/routes/api.v3.batches.ts:178-181
Timestamp: 2026-05-05T09:38:02.512Z
Learning: When reviewing code that catches `ServiceValidationError` in `*.server.ts` files, do not blindly forward `error.status` to HTTP responses, because SVEs may be thrown with non-default statuses (e.g., 400/500) and forwarding them can cause client-visible behavioral regressions (e.g., surfacing 500s to clients). Prefer a safe default response status of `error.status ?? 422`, but only after confirming via the reachable call graph that the caught `ServiceValidationError` instances are expected to carry those non-default statuses; otherwise, normalize to `422` to avoid unexpected client-visible 5xx behavior.
Applied to files:
apps/webapp/app/services/runsReplicationService.server.tsapps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
📚 Learning: 2026-03-29T19:16:28.864Z
Learnt from: nicktrn
Repo: triggerdotdev/trigger.dev PR: 3291
File: apps/webapp/app/v3/featureFlags.ts:53-65
Timestamp: 2026-03-29T19:16:28.864Z
Learning: When reviewing TypeScript code that uses Zod v3, treat `z.coerce.*()` schemas as their direct Zod type (e.g., `z.coerce.boolean()` returns a `ZodBoolean` with `_def.typeName === "ZodBoolean"`) rather than a `ZodEffects`. Only `.preprocess()`, `.refine()`/`.superRefine()`, and `.transform()` are expected to wrap schemas in `ZodEffects`. Therefore, in reviewers’ logic like `getFlagControlType`, do not flag/unblock failures that require unwrapping `ZodEffects` when the input schema is a `z.coerce.*` schema.
Applied to files:
apps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
📚 Learning: 2026-06-09T16:27:26.195Z
Learnt from: myftija
Repo: triggerdotdev/trigger.dev PR: 3878
File: apps/webapp/app/v3/services/computeTemplateCreation.server.ts:0-0
Timestamp: 2026-06-09T16:27:26.195Z
Learning: When working in triggerdotdev/trigger.dev code related to worker-group/region default resolution (e.g., defaultWorkerInstanceGroupId handling used by getGlobalDefaultWorkerGroup, getDefaultWorkerGroupForProject, and RegionsPresenter), do NOT add org-level featureFlags overrides in only one resolution site. That can cause template creation routing/decisions to diverge from actual run routing. If org-level override of the default region/worker group is required, it must be centralized in getGlobalDefaultWorkerGroup so every resolution path remains aligned.
Applied to files:
apps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
📚 Learning: 2026-05-14T08:21:07.614Z
Learnt from: d-cs
Repo: triggerdotdev/trigger.dev PR: 3614
File: apps/webapp/app/v3/mollifier/mollifierGate.server.ts:48-52
Timestamp: 2026-05-14T08:21:07.614Z
Learning: When using Trigger.dev v3 feature flags in the webapp, prefer the existing per-org gating mechanism supported by `flag()` via the `overrides` argument. Pass `Organization.featureFlags` (from `environment.organization.featureFlags`) as the `overrides` value; overrides must take precedence over the global `featureFlag` row. Do not require schema changes or add an `orgId` field to `FlagsOptions` for per-org gating—use the overrides pattern consistently (e.g., in gate flows like `resolveOrgFlag` and any server code that threads `environment.organization.featureFlags` into the gate call).
Applied to files:
apps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts
🪛 ast-grep (0.44.1)
apps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.ts
[warning] 2-2: Importing child_process exposes a command-execution surface; ensure any command/argument built from input is validated, and prefer execFile/spawn with an argument array over exec.
Context: import { fork } from "node:child_process";
Note: [CWE-78] Improper Neutralization of Special Elements used in an OS Command ('OS Command Injection').
(detect-child-process-typescript)
🔇 Additional comments (10)
apps/webapp/test/runsReplicationBenchmark.producer.ts (1)
10-28: LGTM!Also applies to: 121-123, 136-165, 192-192, 203-203
apps/webapp/test/runsReplicationJsonRecoveryBenchmark.test.ts (1)
1-267: LGTM!apps/webapp/app/v3/eventRepository/clickhouseEventRepository.server.ts (2)
360-483: LGTM!
3004-3017: LGTM!apps/webapp/test/clickhouseEventRepositoryJsonRecovery.test.ts (1)
25-123: LGTM!apps/webapp/app/services/runsReplicationService.server.ts (3)
901-905: LGTM!
1129-1256: LGTM!
1607-1630: LGTM!.server-changes/fix-runs-list-batch-drop-on-unparseable-json.md (1)
1-6: LGTM!apps/webapp/test/runsReplicationService.part10.test.ts (1)
18-143: LGTM!
|
Addressed both CodeRabbit nitpicks in the follow-up commit:
|
5ced396 to
f12fa44
Compare
…estable JSON When a run's output/payload or a trace span's attributes contained JSON ClickHouse could not ingest (e.g. an object nested past its depth limit), the analytics insert aborted the whole batch. Because runs and events are written in per-batch inserts, one bad row made every other run or span in that batch go missing from the runs list, traces, and logs, even though the Postgres rows were fine. Both ClickHouse write paths now fall back to retrying the insert with input_format_allow_errors so ClickHouse skips only the rows it can't parse and commits the rest. The healthy insert path is unchanged; a successful row isolation logs at info, only a failed isolation insert is an error.
…tion drop/isolation counters Rename the event repository counter from ingest.flush.rows_isolated to ingest.flush.batches_row_isolated (it counts recovered batches, not rows, and the OTLP->Prometheus suffix would read as a row count otherwise). Add matching runs_replication.batches_row_isolated and runs_replication.batches_dropped OTEL counters so run-replication drops and isolations are alertable in production, not just readable in tests.
… their JSON stripped Extract the Cannot-parse-JSON recovery into one helper shared by the run replication and trace/event writers instead of a near-duplicate copy in each. Recovery now isolates the un-ingestable rows by bisection and re-inserts each with its JSON column emptied, so the row still lands: a run keeps its terminal status and a span keeps its place in the trace, and only the un-ingestable content is dropped. Clean rows land in full; a row that cannot parse even stripped is dropped and counted.
Bound the per-batch isolation work so a burst of un-ingestable rows cannot lag the shared replication stream. Once a batch exceeds the isolation-insert budget the remaining rows are stripped and landed in one insert instead of bisected further. The ceiling is transparent for normal recovery and small run batches (they still isolate precisely); it only bites on a large batch that is heavily poisoned. Adds a recovery_cap_hits counter as the flood signal.
…rt the isolation cap bounds inserts The whole-batch-dropped counter fired whenever a recovered batch stripped no rows but dropped at least one, which mislabels a batch where clean rows landed alongside one unrecoverable row. Gate it on rowsDropped === batchSize (nothing landed) instead. The cap unit test now also asserts the bounded insert count so a broken budget cannot pass unnoticed.
…ison JSON iteratively The ELU monitor called eventLoopUtilization() with no args, which returns the cumulative average since thread start, so each sample smoothed out the spikes the percentile stats are meant to catch; sample the delta between consecutive readings instead. The benchmark producer now builds the deeply nested poison JSON as a string iteratively (and validates the depth) so it does not depend on runtime-specific deep JSON.stringify support.
f12fa44 to
c5f1734
Compare
@trigger.dev/build
trigger.dev
@trigger.dev/core
@trigger.dev/python
@trigger.dev/react-hooks
@trigger.dev/redis-worker
@trigger.dev/rsc
@trigger.dev/schema-to-json
@trigger.dev/sdk
commit: |
db0acab to
0abcce5
Compare
…whole ClickHouse batch When a single run output, trace span, or payload contained JSON ClickHouse could not ingest (for example nesting past its depth limit), the whole insert batch was rejected and those runs and spans silently vanished from the runs list, traces, and logs. Recovery is now per-table: - Runs: follow ClickHouse's failing-row hint to strip just the un-ingestable JSON column(s) so the run still lands with its status, up to a configurable limit (RUN_REPLICATION_MAX_POISON_STRIPS_PER_BATCH, default 1). Past the limit, land the rest with allow_errors and skip the remainder, so recovery cost stays flat on large flushes instead of re-sending the batch per row. - Trace events and payloads: land the batch with allow_errors so the good rows land in one pass and only the un-ingestable rows are skipped. Reading the failing-row hint needs a patch to @clickhouse/client-common, whose error parser otherwise discards the row number from the server response.
0abcce5 to
52a5cba
Compare
Summary
A single run output, trace span, or payload carrying JSON that ClickHouse can't ingest (for example nesting past its depth limit) used to fail the whole insert batch, so unrelated runs and spans silently disappeared from the runs list, traces, and logs. This keeps the rest of the batch and handles the offending row instead of dropping everything around it.
Fix
Recovery is per-table, matched to what each table needs:
task_runs_v2) keep their status. We follow ClickHouse's failing-row hint to strip just the un-ingestable JSON column(s) so the run still lands (its output reads from Postgres on the detail page), up to a configurable limit (RUN_REPLICATION_MAX_POISON_STRIPS_PER_BATCH, default1). Past the limit we stop and land the batch withallow_errorsin a single pass, skipping the remainder. Cost stays a fixed handful of inserts no matter how large or poisoned a flush is.allow_errorsinsert: the good rows land in one pass and only the un-ingestable rows are skipped.Before falling back, a lightweight sanitizer still repairs what it can losslessly (lone UTF-16 surrogates, out-of-range integers) so a repairable row lands in full.
To read the failing-row hint we patch
@clickhouse/client-common: its error parser truncates the server response and discards the(at row N)position, so the patch preserves the full text for the recovery path to read.