feat(run-engine): fair virtual-time scheduling for the concurrency-key dequeue#4367
feat(run-engine): fair virtual-time scheduling for the concurrency-key dequeue#43671stvamp wants to merge 16 commits into
Conversation
|
WalkthroughAdds opt-in CK virtual-time scheduling to the run queue. The feature introduces environment and engine configuration, queue key helpers, Redis Lua commands for fair enqueue, dequeue, and nack handling, virtual-time state with TTL and cleanup behavior, and updated Redis command typings. New integration suites cover ordering, fairness, batching, concurrency, registration, garbage collection, disabled-mode compatibility, and Redis command overhead. Design, rollout, limitation, and research documentation are also included. 🚥 Pre-merge checks | ✅ 3 | ❌ 2❌ Failed checks (2 warnings)
✅ Passed checks (3 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 6
🧹 Nitpick comments (1)
internal-packages/run-engine/src/run-queue/tests/ckVtime.test.ts (1)
11-95: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚖️ Poor tradeoffConsider extracting the duplicated CK-vtime test fixtures.
testOptions,authenticatedEnvDev,makeMessage,variantName, andcreateQueueare near-identically redefined in all three new suites; a shared helper module would keep env/queue config in one place and prevent drift as the vtime option surface evolves.
internal-packages/run-engine/src/run-queue/tests/ckVtime.test.ts#L11-L95: movetestOptions/authenticatedEnvDev/makeMessage/variantName/createQueueinto a shared fixture and import them (keep theVtimeOverrides | nullvariant ofcreateQueuehere).internal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.ts#L22-L94: import the shared fixtures instead of re-declaring; reuse the sharedcreateQueue(...)(keyPrefix +vtimeEnabledvariant).internal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.ts#L29-L94: import the shared fixtures instead of re-declaring.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro Plus
Run ID: e174bf96-1f4a-4e37-bea6-bc9cdcb8d58e
⛔ Files ignored due to path filters (2)
docs/superpowers/references/diagrams/fairness-how-it-works.pngis excluded by!**/*.pngdocs/superpowers/references/diagrams/fairness-problem-and-fix.pngis excluded by!**/*.png
📒 Files selected for processing (19)
.server-changes/2026-07-24-ck-fair-scheduling.mdapps/webapp/app/env.server.tsapps/webapp/app/v3/runEngine.server.tsdocs/superpowers/plans/2026-07-23-ck-virtual-time-scheduling-plan.mddocs/superpowers/references/README.mddocs/superpowers/references/run-queue-fairness-base-queue-findings.mddocs/superpowers/references/run-queue-fairness-caps-vs-scheduling-findings.mddocs/superpowers/references/run-queue-fairness-ck-findings.mddocs/superpowers/references/run-queue-fairness-research.mdinternal-packages/run-engine/src/engine/index.tsinternal-packages/run-engine/src/engine/types.tsinternal-packages/run-engine/src/run-queue/CK_VTIME_KNOWN_LIMITATIONS.mdinternal-packages/run-engine/src/run-queue/index.tsinternal-packages/run-engine/src/run-queue/keyProducer.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsinternal-packages/run-engine/src/run-queue/types.ts
📜 Review details
⏰ Context from checks skipped due to timeout. (18)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (9, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (10, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (7, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (8, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (1, 12)
- GitHub Check: runops-guard / runops-guard
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (4, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (12, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (11, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (2, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (5, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (6, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (3, 12)
- GitHub Check: typecheck / typecheck
- GitHub Check: e2e-webapp / 🧪 E2E Tests: Webapp
- GitHub Check: internal / 🧪 Unit Tests: Internal
- GitHub Check: code-quality / code-quality
- GitHub Check: Analyze (javascript-typescript)
⚠️ CI failures not shown inline (2)
GitHub Actions: 📚 Docs Checks / check-broken-links: feat(run-engine): fair virtual-time scheduling for the concurrency-key dequeue
Conclusion: failure
##[group]Run npx mintlify@4.0.393 broken-links
�[36;1mnpx mintlify@4.0.393 broken-links�[0m
shell: /usr/bin/bash -e {0}
##[endgroup]
Checking for broken links...
Error: Syntax error - Unable to parse superpowers/plans/2026-07-23-ck-virtual-time-scheduling-plan.md - 639:4: Unexpected character `=` (U+003D) before name, expected a character that can start a name, such as a letter, `$`, or `_`
##[error]Process completed with exit code 1.
GitHub Actions: 📚 Docs Checks / 0_check-broken-links.txt: feat(run-engine): fair virtual-time scheduling for the concurrency-key dequeue
Conclusion: failure
##[group]Run npx mintlify@4.0.393 broken-links
�[36;1mnpx mintlify@4.0.393 broken-links�[0m
shell: /usr/bin/bash -e {0}
##[endgroup]
Checking for broken links...
Error: Syntax error - Unable to parse superpowers/plans/2026-07-23-ck-virtual-time-scheduling-plan.md - 639:4: Unexpected character `=` (U+003D) before name, expected a character that can start a name, such as a letter, `$`, or `_`
##[error]Process completed with exit code 1.
🧰 Additional context used
📓 Path-based instructions (11)
**/*.{ts,tsx}
📄 CodeRabbit inference engine (.github/copilot-instructions.md)
**/*.{ts,tsx}: Use types over interfaces for TypeScript
Avoid using enums; prefer string unions or const objects instead
**/*.{ts,tsx}: Prefer static imports over dynamicimport(); use dynamic imports only for unresolvable circular dependencies, genuine performance code splitting, or conditional runtime loading.
Import Trigger.dev tasks from@trigger.dev/sdk; never use@trigger.dev/sdk/v3or deprecatedclient.defineJob.
Add agentcrumbs while writing code using approved namespaces; mark lines with//@Crumbsor blocks with `// `#region` `@crumbs, and strip them before merging.
Files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsapps/webapp/app/v3/runEngine.server.tsinternal-packages/run-engine/src/engine/index.tsinternal-packages/run-engine/src/engine/types.tsapps/webapp/app/env.server.tsinternal-packages/run-engine/src/run-queue/types.tsinternal-packages/run-engine/src/run-queue/keyProducer.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.tsinternal-packages/run-engine/src/run-queue/index.ts
**/*.{ts,tsx,js,jsx}
📄 CodeRabbit inference engine (.github/copilot-instructions.md)
Use function declarations instead of default exports
Files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsapps/webapp/app/v3/runEngine.server.tsinternal-packages/run-engine/src/engine/index.tsinternal-packages/run-engine/src/engine/types.tsapps/webapp/app/env.server.tsinternal-packages/run-engine/src/run-queue/types.tsinternal-packages/run-engine/src/run-queue/keyProducer.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.tsinternal-packages/run-engine/src/run-queue/index.ts
**/*.{test,spec}.{ts,tsx}
📄 CodeRabbit inference engine (.github/copilot-instructions.md)
Use vitest for all tests in the Trigger.dev repository
**/*.{test,spec}.{ts,tsx}: Use Vitest exclusively and never mock dependencies; use Testcontainers for integration dependencies.
Place test files next to the source files they test.
Files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.ts
**/*.ts
📄 CodeRabbit inference engine (.cursor/rules/otel-metrics.mdc)
**/*.ts: When creating or editing OTEL metrics (counters, histograms, gauges), ensure metric attributes have low cardinality by using only enums, booleans, bounded error codes, or bounded shard IDs
Do not use high-cardinality attributes in OTEL metrics such as UUIDs/IDs (envId, userId, runId, projectId, organizationId), unbounded integers (itemCount, batchSize, retryCount), timestamps (createdAt, startTime), or free-form strings (errorMessage, taskName, queueName)
When exporting OTEL metrics via OTLP to Prometheus, be aware that the exporter automatically adds unit suffixes to metric names (e.g., 'my_duration_ms' becomes 'my_duration_ms_milliseconds', 'my_counter' becomes 'my_counter_total'). Account for these transformations when writing Grafana dashboards or Prometheus queries
Files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsapps/webapp/app/v3/runEngine.server.tsinternal-packages/run-engine/src/engine/index.tsinternal-packages/run-engine/src/engine/types.tsapps/webapp/app/env.server.tsinternal-packages/run-engine/src/run-queue/types.tsinternal-packages/run-engine/src/run-queue/keyProducer.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.tsinternal-packages/run-engine/src/run-queue/index.ts
internal-packages/**/*.{ts,tsx}
📄 CodeRabbit inference engine (AGENTS.md)
For internal packages, use
typecheckfor verification and never usebuildas the correctness check.
Files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsinternal-packages/run-engine/src/engine/index.tsinternal-packages/run-engine/src/engine/types.tsinternal-packages/run-engine/src/run-queue/types.tsinternal-packages/run-engine/src/run-queue/keyProducer.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.tsinternal-packages/run-engine/src/run-queue/index.ts
{packages/core,apps/webapp}/**/*.{ts,tsx}
📄 CodeRabbit inference engine (.github/copilot-instructions.md)
Use zod for validation in packages/core and apps/webapp
Files:
apps/webapp/app/v3/runEngine.server.tsapps/webapp/app/env.server.ts
apps/webapp/**/*.{ts,tsx}
📄 CodeRabbit inference engine (.cursor/rules/webapp.mdc)
apps/webapp/**/*.{ts,tsx}: Access environment variables through theenvexport ofenv.server.tsinstead of directly accessingprocess.env
Use subpath exports from@trigger.dev/corepackage instead of importing from the root@trigger.dev/corepathDo not reintroduce the removed v1 execution path;
RunEngineVersion.V1branches may only reject or finalize gracefully so v3 clients receive a clean 4xx, never a 5xx.
Files:
apps/webapp/app/v3/runEngine.server.tsapps/webapp/app/env.server.ts
apps/**/*.{ts,tsx}
📄 CodeRabbit inference engine (AGENTS.md)
For apps, use
typecheckfor verification and never usebuildas the correctness check.
Files:
apps/webapp/app/v3/runEngine.server.tsapps/webapp/app/env.server.ts
apps/webapp/app/**/*.{ts,tsx}
📄 CodeRabbit inference engine (apps/webapp/CLAUDE.md)
apps/webapp/app/**/*.{ts,tsx}: For dashboard changes, visually verify the running Remix app with Chrome DevTools MCP, using snapshots, screenshots, interaction, and console-message checks as appropriate.
UseuseCallbackanduseMemoonly for context provider values, expensive derived data used as a dependency, or stable references required by dependency arrays; do not wrap ordinary event handlers or trivial computations.
Use named constants for sentinel or placeholder values instead of scattering raw string literals across comparisons.
Files:
apps/webapp/app/v3/runEngine.server.tsapps/webapp/app/env.server.ts
apps/webapp/app/**/*.ts
📄 CodeRabbit inference engine (apps/webapp/CLAUDE.md)
apps/webapp/app/**/*.ts: Never userequest.signalto detect client disconnects. UsegetRequestAbortSignal()fromapp/services/httpAsyncStorage.server.ts, which is wired to Express response close events.
Access environment variables through theenvexport fromapp/env.server.ts; never useprocess.envdirectly.
Always use PrismafindFirstinstead offindUnique.
Always use the$transactionhelper from~/db.server, never callprisma.$transactionor$replica.$transactiondirectly. Pass isolation levels as strings, useSerializablefor correctness-critical read-then-write invariants, and guard possibly undefined helper results when a definite value is required.
Files:
apps/webapp/app/v3/runEngine.server.tsapps/webapp/app/env.server.ts
apps/webapp/app/v3/**/*.ts
📄 CodeRabbit inference engine (apps/webapp/CLAUDE.md)
New code must target Run Engine V2 through the singleton in
app/v3/runEngine.server.ts; do not reintroduce V1 execution paths. V1 branches may only reject or finalize gracefully with a clean 4xx.
Files:
apps/webapp/app/v3/runEngine.server.ts
🧠 Learnings (21)
📚 Learning: 2026-03-22T13:26:12.060Z
Learnt from: ericallam
Repo: triggerdotdev/trigger.dev PR: 3244
File: apps/webapp/app/components/code/TextEditor.tsx:81-86
Timestamp: 2026-03-22T13:26:12.060Z
Learning: In the triggerdotdev/trigger.dev codebase, do not flag `navigator.clipboard.writeText(...)` calls for `missing-await`/`unhandled-promise` issues. These clipboard writes are intentionally invoked without `await` and without `catch` handlers across the project; keep that behavior consistent when reviewing TypeScript/TSX files (e.g., usages like in `apps/webapp/app/components/code/TextEditor.tsx`).
Applied to files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsapps/webapp/app/v3/runEngine.server.tsinternal-packages/run-engine/src/engine/index.tsinternal-packages/run-engine/src/engine/types.tsapps/webapp/app/env.server.tsinternal-packages/run-engine/src/run-queue/types.tsinternal-packages/run-engine/src/run-queue/keyProducer.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.tsinternal-packages/run-engine/src/run-queue/index.ts
📚 Learning: 2026-03-22T19:24:14.403Z
Learnt from: matt-aitken
Repo: triggerdotdev/trigger.dev PR: 3187
File: apps/webapp/app/v3/services/alerts/deliverErrorGroupAlert.server.ts:200-204
Timestamp: 2026-03-22T19:24:14.403Z
Learning: In the triggerdotdev/trigger.dev codebase, webhook URLs are not expected to contain embedded credentials/secrets (e.g., fields like `ProjectAlertWebhookProperties` should only hold credential-free webhook endpoints). During code review, if you see logging or inclusion of raw webhook URLs in error messages, do not automatically treat it as a credential-leak/secrets-in-logs issue by default—first verify the URL does not contain embedded credentials (for example, no username/password in the URL, no obvious secret/token query params or fragments). If the URL is credential-free per this project’s conventions, allow the logging.
Applied to files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsapps/webapp/app/v3/runEngine.server.tsinternal-packages/run-engine/src/engine/index.tsinternal-packages/run-engine/src/engine/types.tsapps/webapp/app/env.server.tsinternal-packages/run-engine/src/run-queue/types.tsinternal-packages/run-engine/src/run-queue/keyProducer.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.tsinternal-packages/run-engine/src/run-queue/index.ts
📚 Learning: 2026-05-18T08:21:27.694Z
Learnt from: d-cs
Repo: triggerdotdev/trigger.dev PR: 3632
File: apps/webapp/sentry.server.ts:4-21
Timestamp: 2026-05-18T08:21:27.694Z
Learning: When handling Prisma error P1001 ("Can't reach database server") in TypeScript, don’t assume a single error shape. Prisma can surface P1001 via two different error classes/fields: `PrismaClientKnownRequestError` exposes it as `err.code === "P1001"` (common during mid-query connection drops), while `PrismaClientInitializationError` exposes it as `err.errorCode === "P1001"` (common on client startup failure). Therefore, predicates should use `err.code === "P1001" || err.errorCode === "P1001"`. Do not flag `err.code === "P1001"` as “unreachable/never matches,” as it is expected in production.
Applied to files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsapps/webapp/app/v3/runEngine.server.tsinternal-packages/run-engine/src/engine/index.tsinternal-packages/run-engine/src/engine/types.tsapps/webapp/app/env.server.tsinternal-packages/run-engine/src/run-queue/types.tsinternal-packages/run-engine/src/run-queue/keyProducer.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.tsinternal-packages/run-engine/src/run-queue/index.ts
📚 Learning: 2026-05-18T08:21:27.694Z
Learnt from: d-cs
Repo: triggerdotdev/trigger.dev PR: 3632
File: apps/webapp/sentry.server.ts:4-21
Timestamp: 2026-05-18T08:21:27.694Z
Learning: When handling Prisma errors for P1001 ("Can't reach database server"), do not assume it only appears under a single property name. Prisma may surface P1001 via either `PrismaClientKnownRequestError` (`err.code === "P1001"`, e.g., mid-query connection drops) or `PrismaClientInitializationError` (`err.errorCode === "P1001"`, e.g., client startup connection failure). To reliably detect the condition, check `err.code === "P1001" || err.errorCode === "P1001"`, and avoid review rules that would incorrectly flag `err.code === "P1001"` as unreachable/never-matching.
Applied to files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsapps/webapp/app/v3/runEngine.server.tsinternal-packages/run-engine/src/engine/index.tsinternal-packages/run-engine/src/engine/types.tsapps/webapp/app/env.server.tsinternal-packages/run-engine/src/run-queue/types.tsinternal-packages/run-engine/src/run-queue/keyProducer.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.tsinternal-packages/run-engine/src/run-queue/index.ts
📚 Learning: 2026-06-13T19:53:13.759Z
Learnt from: ericallam
Repo: triggerdotdev/trigger.dev PR: 3937
File: packages/trigger-sdk/skills/realtime-and-frontend/SKILL.md:258-260
Timestamp: 2026-06-13T19:53:13.759Z
Learning: When reviewing code that uses `trigger.dev/react-hooks`’s `useRealtimeRun`, preserve the call signature where the first argument is the full realtime handle object (not `handle.id`). This is intentional to maintain type-safety and is consistent with the official docs; do not suggest changing the first argument from the handle object to `handle.id`.
Applied to files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsapps/webapp/app/v3/runEngine.server.tsinternal-packages/run-engine/src/engine/index.tsinternal-packages/run-engine/src/engine/types.tsapps/webapp/app/env.server.tsinternal-packages/run-engine/src/run-queue/types.tsinternal-packages/run-engine/src/run-queue/keyProducer.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.tsinternal-packages/run-engine/src/run-queue/index.ts
📚 Learning: 2026-06-17T17:13:49.929Z
Learnt from: matt-aitken
Repo: triggerdotdev/trigger.dev PR: 3948
File: apps/webapp/app/routes/_app.orgs.$organizationSlug.projects.$projectParam.env.$envParam.bulk-actions.$bulkActionParam/route.tsx:48-62
Timestamp: 2026-06-17T17:13:49.929Z
Learning: In triggerdotdev/trigger.dev, within `dashboardLoader`/`dashboardAction` (or similar context resolver code) whenever you resolve an organization ID from an organization slug for RBAC/enterprise authorization scope, always read from the primary Prisma client (`prisma`), not `$replica`. Using `$replica` can hit replica-lag and cause the RBAC lookup/authorization to run without the correct org scope (bypassing intended role enforcement). Implement the slug→org lookup with `prisma.organization.findFirst(...)` (or equivalent primary-client query) and add an inline comment documenting why the primary client is required (replica lag could lead to unscoped RBAC checks).
Applied to files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsapps/webapp/app/v3/runEngine.server.tsinternal-packages/run-engine/src/engine/index.tsinternal-packages/run-engine/src/engine/types.tsapps/webapp/app/env.server.tsinternal-packages/run-engine/src/run-queue/types.tsinternal-packages/run-engine/src/run-queue/keyProducer.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.tsinternal-packages/run-engine/src/run-queue/index.ts
📚 Learning: 2026-06-23T13:04:21.413Z
Learnt from: carderne
Repo: triggerdotdev/trigger.dev PR: 4023
File: apps/webapp/app/services/upsertBranch.server.ts:14-18
Timestamp: 2026-06-23T13:04:21.413Z
Learning: In TypeScript, it’s valid to `import { type X }` and then use `typeof X` in a type-only position, e.g. `type Alias = z.infer<typeof X>`. The `type` modifier suppresses the runtime import, but the type checker still has the full exported type so `z.infer<typeof X>` can resolve correctly. In code reviews, don’t flag this as a TypeScript compile error as long as `typeof X` is used in a type context (e.g., with `z.infer`, `type` aliases, generics), not as a runtime value.
Applied to files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsapps/webapp/app/v3/runEngine.server.tsinternal-packages/run-engine/src/engine/index.tsinternal-packages/run-engine/src/engine/types.tsapps/webapp/app/env.server.tsinternal-packages/run-engine/src/run-queue/types.tsinternal-packages/run-engine/src/run-queue/keyProducer.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.tsinternal-packages/run-engine/src/run-queue/index.ts
📚 Learning: 2026-05-18T14:40:02.173Z
Learnt from: ericallam
Repo: triggerdotdev/trigger.dev PR: 3658
File: packages/core/src/v3/realtimeStreams/manager.test.ts:1-147
Timestamp: 2026-05-18T14:40:02.173Z
Learning: In the triggerdotdev/trigger.dev repo, the policy “Never mock anything — use testcontainers instead” should only be enforced for integration tests that interact with real external services (e.g., Redis, Postgres) via actual infrastructure. For unit tests that exercise pure in-memory logic (e.g., cache semantics) it is OK to stub collaborators such as `ApiClient` using Vitest (`vi.fn()`) to assert call counts or control behavior. Do not flag `vi.fn()`-based `ApiClient` stubs in unit tests as violations of the testcontainers policy.
Applied to files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.ts
📚 Learning: 2026-06-04T18:16:35.386Z
Learnt from: nicktrn
Repo: triggerdotdev/trigger.dev PR: 3836
File: apps/supervisor/src/backpressure/backpressureMonitor.ts:3-5
Timestamp: 2026-06-04T18:16:35.386Z
Learning: When reviewing TypeScript in this repo, apply the rule “prefer type aliases over interfaces” only to data/object shapes and union/intersection type modeling. If an interface is being used as a behavioral contract for collaborators to implement (e.g., method-shape interfaces that define required behavior, such as `BackpressureLogger` / `BackpressureSignalSource` in `apps/supervisor/src/backpressure/backpressureMonitor.ts`), keep it as an `interface` and do not flag it as a type-alias-vs-interface violation.
Applied to files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsapps/webapp/app/v3/runEngine.server.tsinternal-packages/run-engine/src/engine/index.tsinternal-packages/run-engine/src/engine/types.tsapps/webapp/app/env.server.tsinternal-packages/run-engine/src/run-queue/types.tsinternal-packages/run-engine/src/run-queue/keyProducer.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.tsinternal-packages/run-engine/src/run-queue/index.ts
📚 Learning: 2026-06-09T17:58:04.699Z
Learnt from: 0ski
Repo: triggerdotdev/trigger.dev PR: 3879
File: apps/webapp/app/models/vercelIntegration.server.ts:619-630
Timestamp: 2026-06-09T17:58:04.699Z
Learning: In this codebase, outbound raw `fetch` calls should typically rely on Node/undici’s default request timeout (about ~300s) rather than adding a per-call `AbortController` + `setTimeout` wrapper inside individual functions (e.g. in files like `apps/webapp/app/models/vercelIntegration.server.ts`). During code review, do not flag the absence of a per-call timeout on a single `fetch` as an issue; if per-call timeouts are needed, they should be implemented via a codebase-wide convention (e.g., a shared fetch wrapper or documented pattern) rather than ad-hoc per-function changes.
Applied to files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsapps/webapp/app/v3/runEngine.server.tsinternal-packages/run-engine/src/engine/index.tsinternal-packages/run-engine/src/engine/types.tsapps/webapp/app/env.server.tsinternal-packages/run-engine/src/run-queue/types.tsinternal-packages/run-engine/src/run-queue/keyProducer.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.tsinternal-packages/run-engine/src/run-queue/index.ts
📚 Learning: 2026-06-16T09:19:47.637Z
Learnt from: d-cs
Repo: triggerdotdev/trigger.dev PR: 3960
File: apps/webapp/test/prismaInfrastructureErrorCapture.test.ts:0-0
Timestamp: 2026-06-16T09:19:47.637Z
Learning: In this repo’s Vitest setup, `vitest.config.ts` uses `globals: true`, so identifiers like `vi`, `describe`, `it`, and `expect` are available as globals in Vitest test files. During code review, do not flag missing `vi`/`describe`/`it`/`expect` imports as a runtime error or correctness issue when they’re used in `*.test.ts/tsx` or `*.spec.ts/tsx` files. Explicit imports are still preferred for consistency, but they’re not required for runtime behavior.
Applied to files:
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.ts
📚 Learning: 2026-05-14T14:54:39.095Z
Learnt from: ericallam
Repo: triggerdotdev/trigger.dev PR: 3545
File: .server-changes/agent-view-sessions.md:10-10
Timestamp: 2026-05-14T14:54:39.095Z
Learning: In the `trigger.dev` repository, do not flag inconsistent dot vs slash notation in route/path strings inside `.server-changes/*.md` files. These markdown files are consumed verbatim into the changelog, so the mixed notation (e.g., `resources.orgs.../runs.$runParam/...`) is intentional and should be preserved as-is.
Applied to files:
.server-changes/2026-07-24-ck-fair-scheduling.md
📚 Learning: 2026-03-29T19:16:28.864Z
Learnt from: nicktrn
Repo: triggerdotdev/trigger.dev PR: 3291
File: apps/webapp/app/v3/featureFlags.ts:53-65
Timestamp: 2026-03-29T19:16:28.864Z
Learning: When reviewing TypeScript code that uses Zod v3, treat `z.coerce.*()` schemas as their direct Zod type (e.g., `z.coerce.boolean()` returns a `ZodBoolean` with `_def.typeName === "ZodBoolean"`) rather than a `ZodEffects`. Only `.preprocess()`, `.refine()`/`.superRefine()`, and `.transform()` are expected to wrap schemas in `ZodEffects`. Therefore, in reviewers’ logic like `getFlagControlType`, do not flag/unblock failures that require unwrapping `ZodEffects` when the input schema is a `z.coerce.*` schema.
Applied to files:
apps/webapp/app/v3/runEngine.server.ts
📚 Learning: 2026-06-09T16:27:26.195Z
Learnt from: myftija
Repo: triggerdotdev/trigger.dev PR: 3878
File: apps/webapp/app/v3/services/computeTemplateCreation.server.ts:0-0
Timestamp: 2026-06-09T16:27:26.195Z
Learning: When working in triggerdotdev/trigger.dev code related to worker-group/region default resolution (e.g., defaultWorkerInstanceGroupId handling used by getGlobalDefaultWorkerGroup, getDefaultWorkerGroupForProject, and RegionsPresenter), do NOT add org-level featureFlags overrides in only one resolution site. That can cause template creation routing/decisions to diverge from actual run routing. If org-level override of the default region/worker group is required, it must be centralized in getGlobalDefaultWorkerGroup so every resolution path remains aligned.
Applied to files:
apps/webapp/app/v3/runEngine.server.ts
📚 Learning: 2026-05-05T09:38:02.512Z
Learnt from: d-cs
Repo: triggerdotdev/trigger.dev PR: 3523
File: apps/webapp/app/routes/api.v3.batches.ts:178-181
Timestamp: 2026-05-05T09:38:02.512Z
Learning: When reviewing code that catches `ServiceValidationError` in `*.server.ts` files, do not blindly forward `error.status` to HTTP responses, because SVEs may be thrown with non-default statuses (e.g., 400/500) and forwarding them can cause client-visible behavioral regressions (e.g., surfacing 500s to clients). Prefer a safe default response status of `error.status ?? 422`, but only after confirming via the reachable call graph that the caught `ServiceValidationError` instances are expected to carry those non-default statuses; otherwise, normalize to `422` to avoid unexpected client-visible 5xx behavior.
Applied to files:
apps/webapp/app/v3/runEngine.server.tsapps/webapp/app/env.server.ts
📚 Learning: 2026-05-12T21:04:05.815Z
Learnt from: ericallam
Repo: triggerdotdev/trigger.dev PR: 3542
File: apps/webapp/app/components/sessions/v1/SessionStatus.tsx:1-3
Timestamp: 2026-05-12T21:04:05.815Z
Learning: In this Remix + TypeScript codebase, do not flag a server/client boundary violation when a file imports only types from a module matching `*.server`.
Specifically, it’s safe to import types using `import type { Foo } from "*.server"` or `import { type Foo } from "*.server"` because TypeScript erases type-only imports at compile time and they emit no JavaScript, so they won’t cross the Remix server/client bundle boundary.
Only raise the boundary concern for value imports (e.g., `import { Foo }` without `type`, or `import Foo`), since those produce JavaScript output.
Applied to files:
apps/webapp/app/v3/runEngine.server.tsapps/webapp/app/env.server.ts
📚 Learning: 2026-06-25T18:21:51.905Z
Learnt from: carderne
Repo: triggerdotdev/trigger.dev PR: 4039
File: apps/webapp/app/routes/invite-revoke.tsx:0-0
Timestamp: 2026-06-25T18:21:51.905Z
Learning: During the Zod v4 migration in the triggerdotdev/trigger.dev webapp, ensure any imports from `conform-to/zod` use the Zod-4 subpath: `conform-to/zod/v4` (e.g., `import { parseWithZod } from "conform-to/zod/v4"`). Do not import from the package root `conform-to/zod`, because it is the Zod 3 implementation and may load Zod-3-only symbols (e.g., `ZodBranded`, `ZodEffects`), which can throw at module load (notably with `zod4.4.3`). This should be enforced across `apps/webapp/**/*` where helpers like `parseWithZod` and `conformZodMessage` are used.
Applied to files:
apps/webapp/app/v3/runEngine.server.tsapps/webapp/app/env.server.ts
📚 Learning: 2026-07-03T17:10:21.498Z
Learnt from: 0ski
Repo: triggerdotdev/trigger.dev PR: 4148
File: apps/webapp/app/models/orgMember.server.ts:149-168
Timestamp: 2026-07-03T17:10:21.498Z
Learning: In triggerdotdev/trigger.dev, `User.email` (Prisma schema: `internal-packages/database/prisma/schema.prisma`) currently does NOT use `citext` and does NOT have a `lower(email)` functional unique index. Therefore, do not introduce Prisma queries like `where: { email: { equals: <value>, mode: "insensitive" } }` (or any case-insensitive lookup) against `User.email`, because it can force sequential scans of the `users` table under load. During review, ensure email is normalized (e.g., lowercased/trimmed) before both writes and subsequent lookups, and if true case-insensitive behavior/uniqueness is required, implement it via a separate app-wide migration (e.g., switch to `citext` and/or add a functional unique index with backfill) rather than bolting it onto individual feature PRs.
Applied to files:
apps/webapp/app/v3/runEngine.server.tsapps/webapp/app/env.server.ts
📚 Learning: 2026-05-14T08:21:07.614Z
Learnt from: d-cs
Repo: triggerdotdev/trigger.dev PR: 3614
File: apps/webapp/app/v3/mollifier/mollifierGate.server.ts:48-52
Timestamp: 2026-05-14T08:21:07.614Z
Learning: When using Trigger.dev v3 feature flags in the webapp, prefer the existing per-org gating mechanism supported by `flag()` via the `overrides` argument. Pass `Organization.featureFlags` (from `environment.organization.featureFlags`) as the `overrides` value; overrides must take precedence over the global `featureFlag` row. Do not require schema changes or add an `orgId` field to `FlagsOptions` for per-org gating—use the overrides pattern consistently (e.g., in gate flows like `resolveOrgFlag` and any server code that threads `environment.organization.featureFlags` into the gate call).
Applied to files:
apps/webapp/app/v3/runEngine.server.ts
📚 Learning: 2026-05-20T17:21:18.543Z
Learnt from: d-cs
Repo: triggerdotdev/trigger.dev PR: 3678
File: apps/webapp/app/entry.server.tsx:0-0
Timestamp: 2026-05-20T17:21:18.543Z
Learning: In env.server.ts (Zod env schema), any environment variable you plan to access via the typed `env` export (e.g., `env.SENTRY_DSN`) must be explicitly declared in the schema. For `SENTRY_DSN`, include `SENTRY_DSN: z.string().optional()`; otherwise switching from `process.env.SENTRY_DSN` to `env.SENTRY_DSN` will fail TypeScript typechecking.
Applied to files:
apps/webapp/app/env.server.ts
📚 Learning: 2026-06-01T11:37:08.569Z
Learnt from: d-cs
Repo: triggerdotdev/trigger.dev PR: 3754
File: apps/webapp/app/env.server.ts:1104-1129
Timestamp: 2026-06-01T11:37:08.569Z
Learning: In apps/*/app/env.server.ts, any new background/periodic worker feature flag should hard-default to "0" (explicit opt-in) rather than inheriting from a parent flag (e.g., avoid defaulting to process.env.TRIGGER_MOLLIFIER_ENABLED ?? "0"). Inheriting can cause the new worker to auto-start on upgrade for deployments that already enabled the parent flag, turning on unexpected background load without an explicit rollout. Each worker component must require its own dedicated env var and default it explicitly to "0" (e.g., TRIGGER_MOLLIFIER_STALE_SWEEP_ENABLED defaults to "0" unless explicitly set to enable that worker).
Applied to files:
apps/webapp/app/env.server.ts
🪛 GitHub Actions: 📚 Docs Checks / 0_check-broken-links.txt
docs/superpowers/plans/2026-07-23-ck-virtual-time-scheduling-plan.md
[error] 639-639: Mintlify broken-links failed: Syntax error while parsing markdown at 639:4. Unexpected character '=' (U+003D) before name; expected a character that can start a name (e.g., a letter, '$', or '_').
🪛 GitHub Actions: 📚 Docs Checks / check-broken-links
docs/superpowers/plans/2026-07-23-ck-virtual-time-scheduling-plan.md
[error] 639-639: Mintlify broken-links check failed. Syntax error: Unable to parse file at 639:4 — Unexpected character '=' (U+003D) before name; expected a character that can start a name (e.g., letter, $, _).
🪛 LanguageTool
docs/superpowers/references/run-queue-fairness-ck-findings.md
[style] ~9-~9: Consider using a shorter alternative to avoid wordiness.
Context: ...e same. A CoDel wrapper on the baseline makes it worse. This was measured by driving the real ...
(MADE_IT_JJR)
[grammar] ~85-~85: Ensure spelling is correct
Context: ...e key per Lua call (maxCount = 1) and rescores ckIndex before each call. Production ...
(QB_NEW_EN_ORTHOGRAPHY_ERROR_IDS_1)
docs/superpowers/references/run-queue-fairness-research.md
[grammar] ~44-~44: Ensure spelling is correct
Context: ...cuts light tenant L's wait to near-zero iff BOTH: 1. Slot availability: sum of cap...
(QB_NEW_EN_ORTHOGRAPHY_ERROR_IDS_1)
[style] ~90-~90: The double modal “needs bounded” is nonstandard (only accepted in certain dialects). Consider “to be bounded”.
Context: ...n one tenant could use all K) and needs bounded, pre-known tenant cardinality. ## Prod...
(NEEDS_FIXED)
[grammar] ~133-~133: Ensure spelling is correct
Context: ...ve-LIFO (newest first), the opposite of stalest-first hoisting.
(QB_NEW_EN_ORTHOGRAPHY_ERROR_IDS_1)
docs/superpowers/plans/2026-07-23-ck-virtual-time-scheduling-plan.md
[style] ~28-~28: ‘exact same’ might be wordy. Consider a shorter alternative.
Context: ...lag off is byte-identical to today: the exact same Lua scripts run and no new Redis keys a...
(EN_WORDINESS_PREMIUM_EXACT_SAME)
[style] ~92-~92: Consider an alternative for the overused word “exactly”.
Context: ...ker variants have older heads (which is exactly where today's age-ordered *3 window f...
(EXACTLY_PRECISELY)
[grammar] ~341-~341: Ensure spelling is correct
Context: ...Dequeue with maxCount: 10 repeatedly (acking between calls to free concurrency). ...
(QB_NEW_EN_ORTHOGRAPHY_ERROR_IDS_1)
[locale-violation] ~356-~356: In American English, ‘afterward’ is the preferred variant. ‘Afterwards’ is more commonly used in British English and other dialects.
Context: ...erved in that first call and its tag afterwards is floor + quantum, not 1. 5. "no s...
(AFTERWARDS_US)
🔇 Additional comments (20)
internal-packages/run-engine/src/run-queue/tests/ckVtime.test.ts (1)
45-74: LGTM!Also applies to: 101-1061
internal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.ts (1)
99-197: LGTM!Also applies to: 199-277, 279-375
internal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.ts (1)
123-206: LGTM!Also applies to: 242-531
internal-packages/run-engine/src/run-queue/tests/keyProducer.test.ts (1)
436-449: LGTM!.server-changes/2026-07-24-ck-fair-scheduling.md (1)
1-6: LGTM!docs/superpowers/references/run-queue-fairness-ck-findings.md (1)
1-122: LGTM!docs/superpowers/references/run-queue-fairness-research.md (1)
1-133: LGTM!internal-packages/run-engine/src/run-queue/CK_VTIME_KNOWN_LIMITATIONS.md (1)
1-71: LGTM!internal-packages/run-engine/src/run-queue/keyProducer.ts (1)
25-26: LGTM!Also applies to: 322-328
internal-packages/run-engine/src/run-queue/types.ts (1)
132-133: LGTM!internal-packages/run-engine/src/run-queue/index.ts (4)
121-135: LGTM!Also applies to: 220-241
1939-2077: LGTM!Also applies to: 2312-2361, 2720-2769
4725-4905: LGTM!
6314-6503: LGTM!apps/webapp/app/env.server.ts (1)
944-949: LGTM!docs/superpowers/references/README.md (1)
1-44: LGTM!docs/superpowers/references/run-queue-fairness-base-queue-findings.md (1)
1-169: LGTM!docs/superpowers/references/run-queue-fairness-caps-vs-scheduling-findings.md (1)
1-199: LGTM!internal-packages/run-engine/src/engine/types.ts (1)
19-19: LGTM!Also applies to: 127-129
internal-packages/run-engine/src/engine/index.ts (1)
237-237: LGTM!
| ckVirtualTimeScheduling: | ||
| env.RUN_ENGINE_CK_VTIME_SCHEDULING_ENABLED | ||
| ? { | ||
| enabled: true, | ||
| quantum: env.RUN_ENGINE_CK_VTIME_QUANTUM, | ||
| scanWindowMultiplier: env.RUN_ENGINE_CK_VTIME_WINDOW_MULTIPLIER, | ||
| stateTtlSeconds: env.RUN_ENGINE_CK_VTIME_STATE_TTL_SECONDS, | ||
| } | ||
| : undefined, |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
Compare the scheduling flag with its enabled value.
The documented environment value is a string defaulting to "0". Direct truthiness makes "0" truthy, enabling virtual-time scheduling by default and violating the rollout contract.
Proposed fix
ckVirtualTimeScheduling:
- env.RUN_ENGINE_CK_VTIME_SCHEDULING_ENABLED
+ env.RUN_ENGINE_CK_VTIME_SCHEDULING_ENABLED === "1"
? {📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| ckVirtualTimeScheduling: | |
| env.RUN_ENGINE_CK_VTIME_SCHEDULING_ENABLED | |
| ? { | |
| enabled: true, | |
| quantum: env.RUN_ENGINE_CK_VTIME_QUANTUM, | |
| scanWindowMultiplier: env.RUN_ENGINE_CK_VTIME_WINDOW_MULTIPLIER, | |
| stateTtlSeconds: env.RUN_ENGINE_CK_VTIME_STATE_TTL_SECONDS, | |
| } | |
| : undefined, | |
| ckVirtualTimeScheduling: | |
| env.RUN_ENGINE_CK_VTIME_SCHEDULING_ENABLED === "1" | |
| ? { | |
| enabled: true, | |
| quantum: env.RUN_ENGINE_CK_VTIME_QUANTUM, | |
| scanWindowMultiplier: env.RUN_ENGINE_CK_VTIME_WINDOW_MULTIPLIER, | |
| stateTtlSeconds: env.RUN_ENGINE_CK_VTIME_STATE_TTL_SECONDS, | |
| } | |
| : undefined, |
| ckVirtualTimeScheduling?: { | ||
| enabled: boolean; | ||
| /** Virtual-time advance per serve (dimensionless). Default 1. */ | ||
| quantum?: number; | ||
| /** Pass-1 candidate window = actualMaxCount * this. Default 3. */ | ||
| scanWindowMultiplier?: number; | ||
| /** EXPIRE applied to ckVtime/ckVtimeFloor on every write. Default 86400. */ | ||
| stateTtlSeconds?: number; | ||
| }; |
There was a problem hiding this comment.
🩺 Stability & Availability | 🟠 Major | ⚡ Quick win
Validate virtual-time configuration before passing it to Lua.
The documented z.coerce.number() settings accept invalid values: non-positive quantum can stop tag progress, fractional window multipliers can produce invalid Redis range bounds, and non-positive TTLs can immediately expire state. Add finite/positive validation and normalize values such as the scan window before enabling the feature.
Also applies to: 718-721
| - ckSybil (the case caps cannot fix): 20 attacker keys x 8 msgs each, all | ||
| older heads, 1 light key x 10 newer. Assert: flag ON mean light wait | ||
| <= 0.7 x flag OFF (spike: 1765 -> 1009), AND light key's first serve | ||
| happens within the first 3 steps (reachability at the floor), AND | ||
| contention-window share: over the steps where >= 2 keys have queued | ||
| backlog, light's served fraction >= 0.5 x its fair share 1/21 (directional, | ||
| per the spike's confounding caveat; wait is the headline). |
There was a problem hiding this comment.
🎯 Functional Correctness | 🔴 Critical | ⚡ Quick win
Fix the invalid Markdown comparison lines.
Mintlify fails at Line 639 because an indented continuation starts with <=; the following >= line has the same parser hazard. Rephrase these as prose or wrap the expressions in inline code.
Proposed documentation fix
- <= 0.7 x flag OFF (spike: 1765 -> 1009), AND light key's first serve
+ The light-key mean wait is no more than 0.7x flag OFF (spike: 1765 -> 1009), AND its first serve
happens within the first 3 steps (reachability at the floor), AND
...
- >= 0.5 x its fair share 1/21 (directional,
+ the light key's served fraction is at least 0.5x its fair share 1/21 (directional,📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| - ckSybil (the case caps cannot fix): 20 attacker keys x 8 msgs each, all | |
| older heads, 1 light key x 10 newer. Assert: flag ON mean light wait | |
| <= 0.7 x flag OFF (spike: 1765 -> 1009), AND light key's first serve | |
| happens within the first 3 steps (reachability at the floor), AND | |
| contention-window share: over the steps where >= 2 keys have queued | |
| backlog, light's served fraction >= 0.5 x its fair share 1/21 (directional, | |
| per the spike's confounding caveat; wait is the headline). | |
| - ckSybil (the case caps cannot fix): 20 attacker keys x 8 msgs each, all | |
| older heads, 1 light key x 10 newer. Assert: flag ON mean light wait | |
| The light-key mean wait is no more than 0.7x flag OFF (spike: 1765 -> 1009), AND its first serve | |
| happens within the first 3 steps (reachability at the floor), AND | |
| contention-window share: over the steps where >= 2 keys have queued | |
| backlog, the light key's served fraction is at least 0.5x its fair share 1/21 (directional, | |
| per the spike's confounding caveat; wait is the headline). |
🧰 Tools
🪛 GitHub Actions: 📚 Docs Checks / 0_check-broken-links.txt
[error] 639-639: Mintlify broken-links failed: Syntax error while parsing markdown at 639:4. Unexpected character '=' (U+003D) before name; expected a character that can start a name (e.g., a letter, '$', or '_').
🪛 GitHub Actions: 📚 Docs Checks / check-broken-links
[error] 639-639: Mintlify broken-links check failed. Syntax error: Unable to parse file at 639:4 — Unexpected character '=' (U+003D) before name; expected a character that can start a name (e.g., letter, $, _).
Source: Pipeline failures
| 17. "op-count budget": using a second plain Redis client, `CONFIG RESETSTAT`, | ||
| run 50 identical dequeue calls flag OFF, snapshot | ||
| `INFO commandstats` total calls; repeat flag ON with identical data. | ||
| Assert `on_total <= off_total + 50 * (6 + 2 * maxCount)` (per call the | ||
| vtime path adds at worst: GET floor, ZRANGE min, ZRANGE window, SET | ||
| floor, EXPIRE, the pass-2 ZRANGEBYSCORE, plus per serve one ZSCORE and | ||
| one ZADD). This pins the per-dequeue overhead the way the caps plan pins | ||
| the fairQueue snapshot cost. |
There was a problem hiding this comment.
🚀 Performance & Scalability | 🟠 Major | ⚡ Quick win
Count skipped candidate attempts in the Redis operation budget.
The budget counts only successful serves, but pass 1 and pass 2 can each call tryServe for up to the scan window of at-cap, future-scheduled, or empty variants. Those attempts still execute Redis operations such as SCARD and ZRANGEBYSCORE, so the asserted upper bound understates worst-case overhead. Include attempted candidates and all per-candidate operations in the budget.
| this.redis.defineCommand("enqueueMessageCkVtimeTracked", { | ||
| numberOfKeys: 17, |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Cross-check declared numberOfKeys against KEYS[n] usage in each new vtime Lua command.
fd -e ts run-queue/index.ts | head -1 | xargs -I{} awk '
/defineCommand\("(enqueueMessageCkVtimeTracked|enqueueMessageWithTtlCkVtimeTracked|dequeueMessagesFromCkQueueVtimeTracked|nackMessageCkVtimeTracked)"/ {cmd=$0}
/numberOfKeys:/ {print cmd; print $0}
' {}
# For each command body, find the highest KEYS[n] referenced:
fd -e ts run-queue/index.ts | head -1 | xargs -I{} grep -noE "KEYS\[[0-9]+\]" {} | sort -t'[' -k2 -n | tail -20Repository: triggerdotdev/trigger.dev
Length of output: 163
🏁 Script executed:
#!/bin/bash
set -euo pipefail
# Locate the exact file and relevant symbols without assuming cwd subtrees.
printf 'Tracked files matching run-queue/index.ts:\n'
git ls-files | rg '(^|/)run-queue/index\.ts$' || true
printf '\nRelevant defineCommand snippets:\n'
while IFS= read -r f; do
echo "--- $f"
rg -n 'defineCommand\("(enqueueMessageCkVtimeTracked|enqueueMessageWithTtlCkVtimeTracked|dequeueMessagesFromCkQueueVtimeTracked|nackMessageCkVtimeTracked)"|numberOfKeys:' "$f" -C 2 || true
done < <(git ls-files | rg '(^|/)run-queue/index\.ts$')
printf '\nAll KEYS references in candidate file(s):\n'
while IFS= read -r f; do
echo "--- $f"
rg -n 'KEYS\[[0-9]+\]' "$f" || true
done < <(git ls-files | rg '(^|/)run-queue/index\.ts$')
printf '\nCall sites of tracked commands:\n'
while IFS= read -r f; do
echo "--- $f"
rg -n 'runCommand\(.*\{[\s\S]{0,1200}numberOfKeys|runCommand\((enqueueMessageCkVtimeTracked|enqueueMessageWithTtlCkVtimeTracked|dequeueMessagesFromCkQueueVtimeTracked|nackMessageCkVtimeTracked)' "$f" || true
done < <(git ls-files | rg '(^|/)run-queue/index\.ts$')Repository: triggerdotdev/trigger.dev
Length of output: 345
🏁 Script executed:
#!/bin/bash
set -euo pipefail
f="internal-packages/run-engine/src/run-queue/index.ts"
printf 'Relevant defineCommand snippets:\n'
rg -n 'defineCommand\("(enqueueMessageCkVtimeTracked|enqueueMessageWithTtlCkVtimeTracked|dequeueMessagesFromCkQueueVtimeTracked|nackMessageCkVtimeTracked)"|numberOfKeys:' "$f" -C 2 || true
printf '\nAll KEYS references in candidate file:\n'
rg -n 'KEYS\[[0-9]+\]' "$f" || true
printf '\nCommand calls/resolver invocations:\n'
rg -n 'runCommand|redis\.(defineCommand|call|callCommand|runCommand)|enqueueMessageCkVtimeTracked|enqueueMessageWithTtlCkVtimeTracked|dequeueMessagesFromCkQueueVtimeTracked|nackMessageCkVtimeTracked' "$f" || trueRepository: triggerdotdev/trigger.dev
Length of output: 50381
🏁 Script executed:
#!/bin/bash
set -euo pipefail
f="internal-packages/run-engine/src/run-queue/index.ts"
python3 - <<'PY' "$f"
import re, sys
from pathlib import Path
path=sys.argv[1]
text=Path(path).read_text()
names=[
"enqueueMessageCkVtimeTracked",
"enqueueMessageWithTtlCkVtimeTracked",
"dequeueMessagesFromCkQueueVtimeTracked",
"nackMessageCkVtimeTracked",
]
for name in names:
m=re.search(rf'defineCommand\("{re.escape(name)}"\s*,\s*\{{(?P<body>.*?)\}},', text, re.S)
if not m:
print(f"{name}: not found")
continue
body=m.group("body")
declared=[int(x) for x in re.findall(r"numberOfKeys:\s*(\d+)", body)]
keys=[int(x) for x in re.findall(r"KEYS\[(\d+)\]", body)]
max_key=max(keys) if keys else 0
call_args=[]
for c in re.finditer(rf"\b{name}\((.*?)\)", text, re.S | re.M):
call=c.group(1)
args=[]
# simple comma split inside call body enough to count position
# capture only first top-level line if it contains arrays args
for line in call.splitlines():
args=re.findall(r'\[([^\]]+)\]|\{(?!\s*.*\}).*?\}', line)
if args:
call_args.append((call, len(args)))
print(f"{name}")
print(f" declared={declared[0] if declared else None} max_KEY_usage={max_key}")
print(f" KEYS refs={sorted(set(keys))}")
print(f" call_arg_count_samples={call_args}")
PY
printf '\nLocks for RedisCommander/defineCommand typing if present:\n'
git ls-files | rg '(^|/)(package\.json|pnpm-lock\.yaml|package-lock\.json|yarn\.lock)$' || true
rg -n 'redis.defineCommand|RedisCommander|defineCommand' package.json pnpm-lock.yaml package-lock.json yarn.lock 2>/dev/null | head -80 || trueRepository: triggerdotdev/trigger.dev
Length of output: 3637
Fix enqueueMessageCkVtimeTracked to include the 18th KEYS reference in numberOfKeys.
The Lua script references KEYS[18] (ckVtimeFloorKey), but numberOfKeys is 17; defineCommand will shift ckVtimeFloorKey into ARGV and the queued envelope will be stored using the wrong Redis key. Increments should use 18 here; the other vtime command declarations match their KEYS usage.
Task 1 of the CK virtual-time (SFQ) scheduling plan. Adds ckVtimeKeyFromQueue and ckVtimeFloorKeyFromQueue (reusing the same base-queue normalisation as ckIndexKeyFromQueue, so :ck:* and :ck:<value> map to one base key), the RunQueueKeyProducer interface signatures, and byte-exact key tests. No runtime behaviour yet.
Task 2. Adds the off-by-default ckVirtualTimeScheduling option (enabled, quantum, scanWindowMultiplier, stateTtlSeconds) to RunQueueOptions and resolves it into private constructor fields for later tasks to read. No behaviour wired yet.
Task 3 (core). Adds dequeueMessagesFromCkQueueVtimeTracked: a new flag-selected Lua command that orders concurrency-key variants by SFQ virtual time. Pass 1 takes candidates from a new :ckVtime ZSET by lowest tag, runs today's per-candidate serve body verbatim (per-key gate, TTL/normal/stale branches, counters, ckIndex rebalance), advances the served variant's tag within the batch, and GCs empty variants from ckVtime too. Pass 2 fills in today's age order (window clamped to >= maxCount*3) so the command is a strict superset of today: work-conserving and mixed-deploy safe. Monotonic floor in :ckVtimeFloor; new variants enter at the floor. The old tracked/untracked scripts are byte-identical and only run when the flag is off. 10 behaviour tests incl. a discriminating vtime-beats-age-order case.
Task 4. Adds enqueueMessageCkVtimeTracked / enqueueMessageWithTtlCkVtimeTracked: the existing tracked enqueue scripts verbatim plus a slow-path ZADD ckVtime NX at the current floor (never rewinds an advanced tag), so a brand-new key is in the fair order from its first enqueue (the sybil fix). Fast path does not register. Old enqueue scripts unchanged; flag-off byte-identical. Tests 10/12 cover registration-at-floor and fast-path purity; test 4 un-seeded.
Task 5. Adds nackMessageCkVtimeTracked: the existing tracked nack script verbatim plus a slow-path ZADD ckVtime NX at the floor after the CK-index rebalance, so a variant a nack revives from GC rejoins the fair order. Old nack script unchanged; flag-off byte-identical. Test 13 covers nack re-registration; test 14 is a 200-op closure-invariant property test (ckIndex subset of ckVtime). Also disables the background master-queue consumers + worker in this file's createQueue helper (matching the RunQueue test convention) so the tests verify the Lua invariant under controlled ops instead of racing a background consumer (was a ~36% flake).
Task 6. New ckVtimeFairness.test.ts drives the real batched Lua (maxCount=10) with flag-ON-vs-OFF ratio assertions over 5 ported spike scenarios, closing the spike's maxCount=1 fidelity gap. ckSkew/ckTrickle: starved-key wait ON <= 0.3x OFF; ckSybil (the case per-key caps cannot fix): light wait ON <= 0.7x OFF + first serve within 3 steps; ckBalanced no-harm; ckHeavyIdle equal drain steps (work conservation). Uses envConcurrencyLimit 1 to force serialization so the ON/OFF contrast is non-vacuous (one-msg-per-variant-per-call would otherwise hide order at limit 4). Reviewer traced each OFF baseline through the real Lua to confirm the assertions genuinely discriminate.
Task 7. New ckVtimeConcurrency.test.ts: two concurrent consumer instances on one base queue serve every message exactly once (no double-serve, no lost message), ckVtime drains empty and the floor never rewinds; concurrent enqueue-during-dequeue never rewinds a tag; and a per-dequeue op-count budget (<= off + 50*(6+2*maxCount)) pins the vtime overhead. Verifies the atomic-single-Lua correctness story.
Task 8. Flag-off test: with ckVirtualTimeScheduling absent, a mixed enqueue/dequeue/nack/ack sequence creates zero *ckVtime* keys and serves in strict head-timestamp (age) order, i.e. today's behaviour. createQueue extended to accept null for an absent option. The whole run-queue suite (16 files, 146 tests) is green with the feature code present and the flag off, proving byte-identical behaviour.
…dark) Task 9. Threads the ckVirtualTimeScheduling option from RunEngineOptions.queue into the RunQueue constructor, and exposes RUN_ENGINE_CK_VTIME_SCHEDULING_ENABLED (+ quantum, window multiplier, state TTL) env vars in the webapp. Off by default at both layers (env default '0' -> undefined -> today's dequeue path). No behaviour change until explicitly enabled.
Task 10 (final). Adds the .server-changes note (ships dark, off by default), fixes three stale/misleading test comments (keyProducer var naming; the un-seeded test 4 comment; the flag-off KEYS-scan comment now correctly credits redisTest flushall), and applies format. No production logic changed.
Three-model blind review (no Critical). Fixes: - Refresh ckVtimeFloor TTL on enqueue/nack, not just dequeue: a dequeue-quiescent but enqueue-active base queue could expire the floor key while ckVtime survived, making a new variant register at 0 and jump the backlog (fairness inversion). - tostring() the vtime tag advance so a fractional quantum/weight is not truncated to an integer by Redis's Lua-number ZADD conversion (silent tag freeze). - Validate the env config: enable flag now uses BoolEnv (so =true/1/yes work, not only '1'); quantum/windowMultiplier/stateTtlSeconds are int().positive() (a 0 TTL errored SET ... EX 0 and stopped all CK dequeues; a 0 multiplier caused a full ZRANGE scan). Constructor clamps as defense-in-depth. - Tests: floor-not-lost regression (H1), large-N (>window) sharding no-starvation, tightened ckBalanced no-harm bound, corrected op-count budget (EXISTS = 7 fixed). - Softened a dangling plan-doc path in a comment. Old Lua command bodies remain byte-identical (edits are in the ...Vtime... commands + env/wiring/tests only).
Captures the whole-branch review findings deliberately not code-fixed (bounded / self-healing / pre-existing): ckVtime tombstone drift on ack/TTL/DLQ/rollback paths and its 24h-TTL / key-delete mitigation; the pre-existing member-name tie-break among equal tags; future-scheduled variants occupying the pass-1 window under retry storms; and the rollout/rollback sequence.
Implementation and testing plan for the recommended run-queue multi-tenant fairness fix: score the concurrency-key dequeue by SFQ virtual time, layered under the concurrency caps. Keeps the three spike findings and the queueing-theory research as references; the throwaway spike harness/bench code is archived on the remote branch chore/fair-queueing-spike and is not carried onto main. The plan adds a parallel :ckVtime ZSET + floor (leaving ckIndex's timestamp domain intact and mixed-deploy-safe), a flag-selected two-pass dequeue command (vtime order then age-order fallback, so it never serves less than today), and a 19-test suite that exercises the real batched maxCount>1 path the spikes could not.
Two diagrams for the PR/design: the problem-and-fix (oldest-first starvation vs virtual-time turns) and the four-step mechanism with guardrails.
check-broken-links parses every .md under docs/ as MDX and can't parse the plain-
markdown plan (code/angle brackets), failing CI. These are engine design docs, not
user documentation, so relocate docs/superpowers/{plans,references} (plan,
findings, research, diagrams) to internal-packages/run-engine/design/ and fix the
internal path references. Also run oxfmt on runEngine.server.ts (missed after the
review fix wave), fixing code-quality.
bc6dc8d to
99eab9d
Compare
@trigger.dev/build
trigger.dev
@trigger.dev/core
@trigger.dev/python
@trigger.dev/react-hooks
@trigger.dev/redis-worker
@trigger.dev/rsc
@trigger.dev/schema-to-json
@trigger.dev/sdk
commit: |
There was a problem hiding this comment.
Actionable comments posted: 3
🧹 Nitpick comments (1)
internal-packages/run-engine/design/plans/2026-07-23-ck-virtual-time-scheduling-plan.md (1)
594-600: 🗄️ Data Integrity & Integration | 🔵 Trivial | ⚡ Quick winAdd an explicit mixed-deployment membership test.
The proposed property test uses only the new registration path, so it cannot validate the documented case where an old producer adds a
ckIndexmember without addingckVtime. Add a scenario that injects an unregisteredckIndexentry or uses an old-path producer, then verifies pass 2 serves and registers it.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro Plus
Run ID: 4350a696-e4f3-4aee-8d5f-f4d0a6a7b062
⛔ Files ignored due to path filters (2)
internal-packages/run-engine/design/references/diagrams/fairness-how-it-works.pngis excluded by!**/*.pnginternal-packages/run-engine/design/references/diagrams/fairness-problem-and-fix.pngis excluded by!**/*.png
📒 Files selected for processing (19)
.server-changes/2026-07-24-ck-fair-scheduling.mdapps/webapp/app/env.server.tsapps/webapp/app/v3/runEngine.server.tsinternal-packages/run-engine/design/plans/2026-07-23-ck-virtual-time-scheduling-plan.mdinternal-packages/run-engine/design/references/README.mdinternal-packages/run-engine/design/references/run-queue-fairness-base-queue-findings.mdinternal-packages/run-engine/design/references/run-queue-fairness-caps-vs-scheduling-findings.mdinternal-packages/run-engine/design/references/run-queue-fairness-ck-findings.mdinternal-packages/run-engine/design/references/run-queue-fairness-research.mdinternal-packages/run-engine/src/engine/index.tsinternal-packages/run-engine/src/engine/types.tsinternal-packages/run-engine/src/run-queue/CK_VTIME_KNOWN_LIMITATIONS.mdinternal-packages/run-engine/src/run-queue/index.tsinternal-packages/run-engine/src/run-queue/keyProducer.tsinternal-packages/run-engine/src/run-queue/tests/ckVtime.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.tsinternal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.tsinternal-packages/run-engine/src/run-queue/tests/keyProducer.test.tsinternal-packages/run-engine/src/run-queue/types.ts
🚧 Files skipped from review as they are similar to previous changes (15)
- .server-changes/2026-07-24-ck-fair-scheduling.md
- internal-packages/run-engine/src/run-queue/tests/keyProducer.test.ts
- internal-packages/run-engine/src/engine/types.ts
- internal-packages/run-engine/design/references/run-queue-fairness-caps-vs-scheduling-findings.md
- internal-packages/run-engine/design/references/README.md
- apps/webapp/app/env.server.ts
- internal-packages/run-engine/design/references/run-queue-fairness-base-queue-findings.md
- internal-packages/run-engine/src/run-queue/types.ts
- internal-packages/run-engine/src/run-queue/tests/ckVtimeConcurrency.test.ts
- internal-packages/run-engine/src/run-queue/CK_VTIME_KNOWN_LIMITATIONS.md
- apps/webapp/app/v3/runEngine.server.ts
- internal-packages/run-engine/src/engine/index.ts
- internal-packages/run-engine/src/run-queue/tests/ckVtime.test.ts
- internal-packages/run-engine/src/run-queue/tests/ckVtimeFairness.test.ts
- internal-packages/run-engine/src/run-queue/index.ts
📜 Review details
⏰ Context from checks skipped due to timeout. (19)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (8, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (5, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (11, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (7, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (9, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (4, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (10, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (12, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (2, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (3, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (6, 12)
- GitHub Check: webapp / 🧪 Unit Tests: Webapp (1, 12)
- GitHub Check: e2e-webapp / 🧪 E2E Tests: Webapp
- GitHub Check: internal / 🧪 Unit Tests: Internal
- GitHub Check: runops-guard / runops-guard
- GitHub Check: typecheck / typecheck
- GitHub Check: code-quality / code-quality
- GitHub Check: Analyze (javascript-typescript)
- GitHub Check: Build and publish previews
🧰 Additional context used
📓 Path-based instructions (4)
**/*.{ts,tsx}
📄 CodeRabbit inference engine (.github/copilot-instructions.md)
**/*.{ts,tsx}: Use types over interfaces for TypeScript
Avoid using enums; prefer string unions or const objects instead
**/*.{ts,tsx}: Prefer static imports over dynamicimport(); use dynamic imports only for unresolvable circular dependencies, genuine performance code splitting, or conditional runtime loading.
Import Trigger.dev tasks from@trigger.dev/sdk; never use@trigger.dev/sdk/v3or deprecatedclient.defineJob.
Add agentcrumbs while writing code using approved namespaces; mark lines with//@Crumbsor blocks with `// `#region` `@crumbs, and strip them before merging.
Files:
internal-packages/run-engine/src/run-queue/keyProducer.ts
**/*.{ts,tsx,js,jsx}
📄 CodeRabbit inference engine (.github/copilot-instructions.md)
Use function declarations instead of default exports
Files:
internal-packages/run-engine/src/run-queue/keyProducer.ts
**/*.ts
📄 CodeRabbit inference engine (.cursor/rules/otel-metrics.mdc)
**/*.ts: When creating or editing OTEL metrics (counters, histograms, gauges), ensure metric attributes have low cardinality by using only enums, booleans, bounded error codes, or bounded shard IDs
Do not use high-cardinality attributes in OTEL metrics such as UUIDs/IDs (envId, userId, runId, projectId, organizationId), unbounded integers (itemCount, batchSize, retryCount), timestamps (createdAt, startTime), or free-form strings (errorMessage, taskName, queueName)
When exporting OTEL metrics via OTLP to Prometheus, be aware that the exporter automatically adds unit suffixes to metric names (e.g., 'my_duration_ms' becomes 'my_duration_ms_milliseconds', 'my_counter' becomes 'my_counter_total'). Account for these transformations when writing Grafana dashboards or Prometheus queries
Files:
internal-packages/run-engine/src/run-queue/keyProducer.ts
internal-packages/**/*.{ts,tsx}
📄 CodeRabbit inference engine (AGENTS.md)
For internal packages, use
typecheckfor verification and never usebuildas the correctness check.
Files:
internal-packages/run-engine/src/run-queue/keyProducer.ts
🧠 Learnings (9)
📚 Learning: 2026-03-22T13:26:12.060Z
Learnt from: ericallam
Repo: triggerdotdev/trigger.dev PR: 3244
File: apps/webapp/app/components/code/TextEditor.tsx:81-86
Timestamp: 2026-03-22T13:26:12.060Z
Learning: In the triggerdotdev/trigger.dev codebase, do not flag `navigator.clipboard.writeText(...)` calls for `missing-await`/`unhandled-promise` issues. These clipboard writes are intentionally invoked without `await` and without `catch` handlers across the project; keep that behavior consistent when reviewing TypeScript/TSX files (e.g., usages like in `apps/webapp/app/components/code/TextEditor.tsx`).
Applied to files:
internal-packages/run-engine/src/run-queue/keyProducer.ts
📚 Learning: 2026-03-22T19:24:14.403Z
Learnt from: matt-aitken
Repo: triggerdotdev/trigger.dev PR: 3187
File: apps/webapp/app/v3/services/alerts/deliverErrorGroupAlert.server.ts:200-204
Timestamp: 2026-03-22T19:24:14.403Z
Learning: In the triggerdotdev/trigger.dev codebase, webhook URLs are not expected to contain embedded credentials/secrets (e.g., fields like `ProjectAlertWebhookProperties` should only hold credential-free webhook endpoints). During code review, if you see logging or inclusion of raw webhook URLs in error messages, do not automatically treat it as a credential-leak/secrets-in-logs issue by default—first verify the URL does not contain embedded credentials (for example, no username/password in the URL, no obvious secret/token query params or fragments). If the URL is credential-free per this project’s conventions, allow the logging.
Applied to files:
internal-packages/run-engine/src/run-queue/keyProducer.ts
📚 Learning: 2026-05-18T08:21:27.694Z
Learnt from: d-cs
Repo: triggerdotdev/trigger.dev PR: 3632
File: apps/webapp/sentry.server.ts:4-21
Timestamp: 2026-05-18T08:21:27.694Z
Learning: When handling Prisma error P1001 ("Can't reach database server") in TypeScript, don’t assume a single error shape. Prisma can surface P1001 via two different error classes/fields: `PrismaClientKnownRequestError` exposes it as `err.code === "P1001"` (common during mid-query connection drops), while `PrismaClientInitializationError` exposes it as `err.errorCode === "P1001"` (common on client startup failure). Therefore, predicates should use `err.code === "P1001" || err.errorCode === "P1001"`. Do not flag `err.code === "P1001"` as “unreachable/never matches,” as it is expected in production.
Applied to files:
internal-packages/run-engine/src/run-queue/keyProducer.ts
📚 Learning: 2026-05-18T08:21:27.694Z
Learnt from: d-cs
Repo: triggerdotdev/trigger.dev PR: 3632
File: apps/webapp/sentry.server.ts:4-21
Timestamp: 2026-05-18T08:21:27.694Z
Learning: When handling Prisma errors for P1001 ("Can't reach database server"), do not assume it only appears under a single property name. Prisma may surface P1001 via either `PrismaClientKnownRequestError` (`err.code === "P1001"`, e.g., mid-query connection drops) or `PrismaClientInitializationError` (`err.errorCode === "P1001"`, e.g., client startup connection failure). To reliably detect the condition, check `err.code === "P1001" || err.errorCode === "P1001"`, and avoid review rules that would incorrectly flag `err.code === "P1001"` as unreachable/never-matching.
Applied to files:
internal-packages/run-engine/src/run-queue/keyProducer.ts
📚 Learning: 2026-06-13T19:53:13.759Z
Learnt from: ericallam
Repo: triggerdotdev/trigger.dev PR: 3937
File: packages/trigger-sdk/skills/realtime-and-frontend/SKILL.md:258-260
Timestamp: 2026-06-13T19:53:13.759Z
Learning: When reviewing code that uses `trigger.dev/react-hooks`’s `useRealtimeRun`, preserve the call signature where the first argument is the full realtime handle object (not `handle.id`). This is intentional to maintain type-safety and is consistent with the official docs; do not suggest changing the first argument from the handle object to `handle.id`.
Applied to files:
internal-packages/run-engine/src/run-queue/keyProducer.ts
📚 Learning: 2026-06-17T17:13:49.929Z
Learnt from: matt-aitken
Repo: triggerdotdev/trigger.dev PR: 3948
File: apps/webapp/app/routes/_app.orgs.$organizationSlug.projects.$projectParam.env.$envParam.bulk-actions.$bulkActionParam/route.tsx:48-62
Timestamp: 2026-06-17T17:13:49.929Z
Learning: In triggerdotdev/trigger.dev, within `dashboardLoader`/`dashboardAction` (or similar context resolver code) whenever you resolve an organization ID from an organization slug for RBAC/enterprise authorization scope, always read from the primary Prisma client (`prisma`), not `$replica`. Using `$replica` can hit replica-lag and cause the RBAC lookup/authorization to run without the correct org scope (bypassing intended role enforcement). Implement the slug→org lookup with `prisma.organization.findFirst(...)` (or equivalent primary-client query) and add an inline comment documenting why the primary client is required (replica lag could lead to unscoped RBAC checks).
Applied to files:
internal-packages/run-engine/src/run-queue/keyProducer.ts
📚 Learning: 2026-06-23T13:04:21.413Z
Learnt from: carderne
Repo: triggerdotdev/trigger.dev PR: 4023
File: apps/webapp/app/services/upsertBranch.server.ts:14-18
Timestamp: 2026-06-23T13:04:21.413Z
Learning: In TypeScript, it’s valid to `import { type X }` and then use `typeof X` in a type-only position, e.g. `type Alias = z.infer<typeof X>`. The `type` modifier suppresses the runtime import, but the type checker still has the full exported type so `z.infer<typeof X>` can resolve correctly. In code reviews, don’t flag this as a TypeScript compile error as long as `typeof X` is used in a type context (e.g., with `z.infer`, `type` aliases, generics), not as a runtime value.
Applied to files:
internal-packages/run-engine/src/run-queue/keyProducer.ts
📚 Learning: 2026-06-04T18:16:35.386Z
Learnt from: nicktrn
Repo: triggerdotdev/trigger.dev PR: 3836
File: apps/supervisor/src/backpressure/backpressureMonitor.ts:3-5
Timestamp: 2026-06-04T18:16:35.386Z
Learning: When reviewing TypeScript in this repo, apply the rule “prefer type aliases over interfaces” only to data/object shapes and union/intersection type modeling. If an interface is being used as a behavioral contract for collaborators to implement (e.g., method-shape interfaces that define required behavior, such as `BackpressureLogger` / `BackpressureSignalSource` in `apps/supervisor/src/backpressure/backpressureMonitor.ts`), keep it as an `interface` and do not flag it as a type-alias-vs-interface violation.
Applied to files:
internal-packages/run-engine/src/run-queue/keyProducer.ts
📚 Learning: 2026-06-09T17:58:04.699Z
Learnt from: 0ski
Repo: triggerdotdev/trigger.dev PR: 3879
File: apps/webapp/app/models/vercelIntegration.server.ts:619-630
Timestamp: 2026-06-09T17:58:04.699Z
Learning: In this codebase, outbound raw `fetch` calls should typically rely on Node/undici’s default request timeout (about ~300s) rather than adding a per-call `AbortController` + `setTimeout` wrapper inside individual functions (e.g. in files like `apps/webapp/app/models/vercelIntegration.server.ts`). During code review, do not flag the absence of a per-call timeout on a single `fetch` as an issue; if per-call timeouts are needed, they should be implemented via a codebase-wide convention (e.g., a shared fetch wrapper or documented pattern) rather than ad-hoc per-function changes.
Applied to files:
internal-packages/run-engine/src/run-queue/keyProducer.ts
🪛 LanguageTool
internal-packages/run-engine/design/plans/2026-07-23-ck-virtual-time-scheduling-plan.md
[style] ~28-~28: ‘exact same’ might be wordy. Consider a shorter alternative.
Context: ...lag off is byte-identical to today: the exact same Lua scripts run and no new Redis keys a...
(EN_WORDINESS_PREMIUM_EXACT_SAME)
[style] ~92-~92: Consider an alternative for the overused word “exactly”.
Context: ...ker variants have older heads (which is exactly where today's age-ordered *3 window f...
(EXACTLY_PRECISELY)
[grammar] ~341-~341: Ensure spelling is correct
Context: ...Dequeue with maxCount: 10 repeatedly (acking between calls to free concurrency). ...
(QB_NEW_EN_ORTHOGRAPHY_ERROR_IDS_1)
[locale-violation] ~356-~356: In American English, ‘afterward’ is the preferred variant. ‘Afterwards’ is more commonly used in British English and other dialects.
Context: ...erved in that first call and its tag afterwards is floor + quantum, not 1. 5. "no s...
(AFTERWARDS_US)
internal-packages/run-engine/design/references/run-queue-fairness-ck-findings.md
[style] ~9-~9: Consider using a shorter alternative to avoid wordiness.
Context: ...e same. A CoDel wrapper on the baseline makes it worse. This was measured by driving the real ...
(MADE_IT_JJR)
[grammar] ~85-~85: Ensure spelling is correct
Context: ...e key per Lua call (maxCount = 1) and rescores ckIndex before each call. Production ...
(QB_NEW_EN_ORTHOGRAPHY_ERROR_IDS_1)
internal-packages/run-engine/design/references/run-queue-fairness-research.md
[grammar] ~44-~44: Ensure spelling is correct
Context: ...cuts light tenant L's wait to near-zero iff BOTH: 1. Slot availability: sum of cap...
(QB_NEW_EN_ORTHOGRAPHY_ERROR_IDS_1)
[style] ~90-~90: The double modal “needs bounded” is nonstandard (only accepted in certain dialects). Consider “to be bounded”.
Context: ...n one tenant could use all K) and needs bounded, pre-known tenant cardinality. ## Prod...
(NEEDS_FIXED)
[grammar] ~133-~133: Ensure spelling is correct
Context: ...ve-LIFO (newest first), the opposite of stalest-first hoisting.
(QB_NEW_EN_ORTHOGRAPHY_ERROR_IDS_1)
🔇 Additional comments (1)
internal-packages/run-engine/src/run-queue/keyProducer.ts (1)
25-26: LGTM!Also applies to: 322-329
| `internal-packages/run-engine/design/references/run-queue-fairness-ck-findings.md`, | ||
| `.../run-queue-fairness-caps-vs-scheduling-findings.md`, and | ||
| `.../run-queue-fairness-research.md`. |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
Resolve the throwaway-document contradiction.
The plan retains a document that explicitly requires deletion before merge. Remove the spike artifact or promote it into durable documentation and update its header.
internal-packages/run-engine/design/plans/2026-07-23-ck-virtual-time-scheduling-plan.md#L7-L9: stop listing the throwaway findings file as a retained reference unless it is promoted.internal-packages/run-engine/design/references/run-queue-fairness-ck-findings.md#L3-L3: delete the file or remove the throwaway-only merge restriction after rewriting it as durable documentation.
📍 Affects 2 files
internal-packages/run-engine/design/plans/2026-07-23-ck-virtual-time-scheduling-plan.md#L7-L9(this comment)internal-packages/run-engine/design/references/run-queue-fairness-ck-findings.md#L3-L3
| - `.server-changes/2026-XX-XX-ck-fair-scheduling.md` (at PR time; note it | ||
| ships dark) |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
Replace the placeholder changeset path.
The plan still references .server-changes/2026-XX-XX-ck-fair-scheduling.md, while the dated file is .server-changes/2026-07-24-ck-fair-scheduling.md. Leaving the placeholder makes the implementation plan stale.
Also applies to: 755-758
| 16. "concurrent enqueue during dequeue cannot rewind a tag": interleave | ||
| enqueues on a hot key with dequeue batches; after each round assert | ||
| `ZSCORE ckVtime <hot>` is non-decreasing (NX registration + advance-only | ||
| writes). |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Do not assert tag monotonicity across GC re-entry.
The plan explicitly resets a drained variant to the floor when it re-registers, so its ZSCORE can decrease after ZREM and later enqueue. Scope this assertion to periods where the member remains registered, or allow the documented GC/re-entry reset.
Part of #2617. Off by default.
When many concurrency-key variants share one task queue, the dequeue serves the oldest waiting run first, so one key's large backlog is served to exhaustion while keys queued behind it wait for the whole pile to drain. This adds an opt-in fair order: each key gets a virtual clock, the dequeue serves the smallest clock and advances it, so keys take turns instead of one pile draining. With the flag off, the existing scripts run unchanged.
How it works
:ckVtimeZSET (the virtual clocks) and a monotonic floor.ckIndexkeeps its head-timestamp domain, so time-eligibility, master-queue rebalancing, and every other writer stay untouched (this is what makes it mixed-deploy safe).Testing
Full run-queue suite is green with the flag off (no regression). Fairness is proven on the real batched dequeue path (not just one message per call), plus multi-consumer exactly-once, a per-dequeue op-count budget, and behaviour tests for the floor, tag advance, GC, and registration.
Rollout
Off by default behind
RUN_ENGINE_CK_VTIME_SCHEDULING_ENABLED. Enable on a staging cell, then production; rollback is flipping the flag off (leftover state expires within a day). During a rolling deploy, old instances serve in age order and are folded in by pass 2, so nothing is lost and no run is served twice. Every mutation is a single atomic Lua script and a base queue's keys share one hash slot, so this holds on a single Redis and on Redis Cluster alike.Known limitations
Bounded, self-healing, or pre-existing edges are documented in
internal-packages/run-engine/src/run-queue/CK_VTIME_KNOWN_LIMITATIONS.md(worth reading before enabling). The design and testing plan, the queueing-theory research behind the approach, and the diagrams above live underinternal-packages/run-engine/design/.