Skip to content

Feature Request: sqlmesh lint loads the full project context even when scoped to specific models with --model #5907

Description

@fresioAS

Summary

sqlmesh lint --model <mymodel> loads the entire project context, parses all model files, resolves all schemas, and validates all model definitions before a single lint rule runs. --model only scopes the linting loop — not the context initialization that precedes it. Cost scales with total project size, not with the number of models being linted.

Where this happens

1. lint always triggers a full context load

lint is not in SKIP_LOAD_COMMANDS, so GenericContext.__init__ always calls self.load() before any linting happens:
sqlmesh/cli/main.py ~L31–42:

SKIP_LOAD_COMMANDS = (
    "clean", "create_external_models", "destroy", "environments",
    "invalidate", "janitor", "migrate", "rollback", "run", "table_name",
)
# "lint" is not here

sqlmesh/core/context.py ~L473:

if load:
    self.load()

2. load() parses and validates the entire project

load() parses every model, macro, audit, test, and metric, then runs schema resolution and definition validation on every model:
sqlmesh/core/context.py ~L719–728:

update_model_schemas(
    self.dag,
    models=self._models,
    cache_dir=self.cache_dir,
)
models = self.models.values()
for model in models:
    model.validate_definition()

update_model_schemas() traverses the full project DAG to propagate column types/schema info — needed for rules like ambiguousorinvalidcolumn, but run for all models regardless of what's being linted.

3. lint_models() scopes the loop, but not the context

The CLI lint command calls obj.lint_models(models) on the already fully-loaded context. The loop itself is scoped:
sqlmesh/core/context.py ~L3209–3219:

model_list = (
    list(self.get_model(model, raise_if_missing=True) for model in models)
    if models
    else self.models.values()
)
for model in model_list:
    if linter := self._linters.get(model.project):
        lint_violation, violations = (
            linter.lint_model(model, self, console=self.console) or found_error
        )

...but the full context (self) is passed into every rule invocation.

4. Every rule receives the full context

RuleSet.check_model() instantiates each rule with the full context, regardless of whether the rule needs it:
sqlmesh/core/linter/definition.py ~L106–116:

def check_model(self, model: Model, context: GenericContext) -> t.List[RuleViolation]:
    for rule in self._underlying.values():
        violation = rule(context).check_model(model)
        ...

5. Some rules genuinely need cross-model context — but not the whole project

  • NoMissingUnitTest checks self.context.models_with_tests (full test index):
    sqlmesh/core/linter/rules/builtin.py ~L140.
  • NoMissingExternalModels calls self.context.get_model(depends_on_model) per dependency:
    sqlmesh/core/linter/rules/builtin.py ~L160–163.

Both only need the target model's dependencies, not the entire project.

6. Single-model rules pay the same cost as cross-model rules

Rules that only inspect the model itself (NoSelectStar, NoMissingAudits, NoAmbiguousProjections) still receive the full context even though they never use it — there's no mechanism to distinguish rules that need cross-model context from those that don't.

What --model actually does (and doesn't do)

It limits which models are iterated in the linting loop. It does not prevent: loading all model files from disk, running update_model_schemas() on the full project DAG, calling model.validate_definition() on all models, or passing the full context to every rule.

Impact

  • Large projects: sqlmesh lint --model my_model in a project with thousands of models still parses every file and resolves every schema. Overhead scales with project size, not selection size.
  • Iteration loop: sqlmesh lint is meant to be faster than sqlmesh plan for quick iteration; a full context load erodes that advantage.
  • Schema resolution is full-project: update_model_schemas() traverses the entire DAG even for single-model rules like ambiguousorinvalidcolumn, when only the target model's upstream subgraph is needed.
  • Cross-model rules need only a bounded subgraph: e.g. NoMissingExternalModels only needs the target model's direct dependencies, not the full model registry.

Suggested direction

  1. Lazy/partial context loading for lint — when --model is specified, load only the target model(s) and their transitive upstream dependencies, deferring parsing of unrelated models.
  2. Scoped schema resolution — allow update_model_schemas() to run on a subgraph instead of always the full project DAG.
  3. Rule context interface — distinguish rules that need full cross-model context from those that only need the model itself, so the latter can run without a full context load.
  4. Scoped cross-model rules — satisfy rules like NoMissingExternalModels/NoMissingUnitTest with a partial context bounded to the target model's dependencies/tests, rather than the full project load.

References

  • CLI lint command and SKIP_LOAD_COMMANDS: sqlmesh/cli/main.py ~L31–42, ~L1168–1183
  • Unconditional self.load(): sqlmesh/core/context.py ~L473–474
  • load() full parse, schema update, validation: sqlmesh/core/context.py ~L629–728
  • lint_models() scoped loop, full context passed: sqlmesh/core/context.py ~L3209–3230
  • RuleSet.check_model() full context injected into every rule: sqlmesh/core/linter/definition.py ~L106–116
  • NoMissingUnitTest uses context.models_with_tests: sqlmesh/core/linter/rules/builtin.py ~L140
  • NoMissingExternalModels uses context.get_model() per dependency: sqlmesh/core/linter/rules/builtin.py ~L160–163

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions