Skip to content

fix(ci): skip loaded lines calculation on iterations after the first#17790

Open
hebaalazzeh wants to merge 4 commits into
mainfrom
perf/import-profiler-skip-line-count
Open

fix(ci): skip loaded lines calculation on iterations after the first#17790
hebaalazzeh wants to merge 4 commits into
mainfrom
perf/import-profiler-skip-line-count

Conversation

@hebaalazzeh

@hebaalazzeh hebaalazzeh commented Jul 20, 2026

Copy link
Copy Markdown
Contributor

This PR optimizes the pure execution latency and reduces redundant disk I/O of the import profiler script by calculating the deterministic codebase statistics (loaded lines count) only on the first iteration.

Problem

For each package profiled, the script executes 11 iterations for the baseline commit and 11 iterations for the PR commit (22 iterations total).
On every single iteration, the worker process scans and counts the lines of every imported .py file (across standard library and third-party dependencies like grpc, protobuf, google-api-core, etc.). This causes redundant and expensive disk I/O.

Solution

  • First Iteration Only: The first worker run profiles and counts the loaded line count as a baseline.
  • Skip Subsequent Counting: For iterations 2-11 (i > 0), the master passes --skip-line-count to the worker process. The worker skips the file-reading loop entirely and returns -1 for loaded_lines.
  • Improved Logging: Worker process stderr is now forwarded to the parent master process so that warnings are visible in GHA logs.

Benchmarks (Timing the line-counting loop itself)

@hebaalazzeh hebaalazzeh self-assigned this Jul 20, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a --skip-line-count option to the import profiler script to optimize execution by skipping line counting on iterations after the first. A review comment highlights an issue where setting loaded_lines to -1 on skipped iterations will corrupt downstream metrics and reporting, and suggests restoring the baseline value instead.

Comment thread scripts/import_profiler/profiler.py Outdated
@hebaalazzeh
hebaalazzeh force-pushed the perf/import-profiler-skip-line-count branch from c06d421 to a004db7 Compare July 20, 2026 21:49
@hebaalazzeh

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a --skip-line-count optimization to the import profiler, allowing worker processes to skip counting loaded lines on iterations after the first one to improve performance. Feedback on these changes suggests suppressing the internal --skip-line-count flag from the CLI help output to prevent user confusion, and removing a redundant, now-dead warning check for non-deterministic loaded lines on subsequent iterations.

Comment thread scripts/import_profiler/profiler.py Outdated
Comment thread scripts/import_profiler/profiler.py Outdated
@hebaalazzeh
hebaalazzeh force-pushed the perf/import-profiler-skip-line-count branch from a004db7 to 086c03b Compare July 20, 2026 21:54
@hebaalazzeh
hebaalazzeh marked this pull request as ready for review July 20, 2026 21:58
@hebaalazzeh
hebaalazzeh requested a review from a team as a code owner July 20, 2026 21:58
return 0.0

def run_worker(target_module):
def run_worker(target_module, skip_line_count=False):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

where / when do we set skip_line_count to true?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the master process, we append "--skip-line-count" to the spawned worker command line for all iterations after the first one (i.e., i > 0):

if i > 0:
cmd += ["--skip-line-count"]

The worker process then parses this CLI argument and passes it to run_worker:

if args.worker:
run_worker(target_module, skip_line_count=args.skip_line_count)

@hebaalazzeh
hebaalazzeh requested review from a team as code owners July 20, 2026 22:13
@hebaalazzeh
hebaalazzeh removed request for a team July 21, 2026 00:11

@parthea parthea left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please can you add tests for scripts/import_profiler/profiler.py?

We can run the test with pytest test_profiler.py --cov to also check code coverage.

Gemini shared this test file which has ~50% coverage. These tests are mainly for code coverage so we may not need all of the tests. The purpose is to signal potential regressions on subsequent changes.

import csv
import json
import sys
import pytest
from unittest.mock import patch, MagicMock, mock_open

# Import core profiler objects
import profiler
from profiler import (
    run_worker,
    _run_worker_and_parse,
    run_master,
    NO_CPU_PINNING
)

# Dynamically resolve helpers to avoid ImportErrors if defined inside __main__
get_rss_mb_helper = getattr(profiler, "get_rss_mb", None)
find_module_helper = getattr(profiler, "find_module_from_package", None)

# =====================================================================
# 1. UTILITY FUNCTIONS TESTS
# =====================================================================

def test_get_rss_mb():
    """Verifies get_rss_mb returns a float representing megabytes if available."""
    if get_rss_mb_helper is None:
        pytest.skip("get_rss_mb is not exported at the module level.")
    rss = get_rss_mb_helper()
    assert isinstance(rss, float)
    assert rss >= 0.0


@patch('os.walk')
@patch('os.remove')
@patch('shutil.rmtree')
def test_clean_bytecode(mock_rmtree, mock_remove, mock_walk):
    """Verifies clean_bytecode successfully cleans bytecode files & caches."""
    clean_bytecode_helper = getattr(profiler, "clean_bytecode", None)
    if clean_bytecode_helper is None:
        pytest.skip("clean_bytecode is not exported at the module level.")
        
    mock_walk.return_value = [
        ('/test_dir', ['__pycache__'], ['test.pyc', 'test.py'])
    ]
    clean_bytecode_helper()
    assert mock_remove.called or mock_rmtree.called


def test_find_module_from_package_resolves():
    """Verifies resolving of packages to modules if helper is exported."""
    if find_module_helper is None:
        pytest.skip("find_module_from_package is not exported at the module level.")
        
    with patch('importlib.metadata.packages_distributions', return_value={"google-cloud-storage": ["google.cloud.storage"]}):
        res = find_module_helper("google-cloud-storage")
        assert res == "google.cloud.storage"


def test_find_module_from_package_fallback():
    """Verifies fallback transforms work correctly if helper is exported."""
    if find_module_helper is None:
        pytest.skip("find_module_from_package is not exported at the module level.")
    res = find_module_helper("google-cloud-ndb")
    assert res == "google_cloud_ndb"


@patch('subprocess.run')
def test_run_trace(mock_run):
    """Verifies that run_trace executes python with -X importtime and writes logs."""
    run_trace_helper = getattr(profiler, "run_trace", None)
    if run_trace_helper is None:
        pytest.skip("run_trace is not exported at the module level.")
        
    # Mock subprocess output
    mock_run.return_value = MagicMock(stdout="", stderr="importtime: dummy trace output", returncode=0)
    
    with patch('builtins.open', mock_open()) as mock_file, \
         patch('builtins.print'):
        run_trace_helper("math")
        
        # Verify that subprocess was invoked with -X importtime
        called_cmd = mock_run.call_args[0][0]
        assert "-X" in called_cmd
        assert "importtime" in called_cmd
        
        # Verify that it attempted to write the trace log file
        assert mock_file.called


# =====================================================================
# 2. WORKER TESTS (run_worker)
# =====================================================================

def test_run_worker_with_skip_line_count(capsys):
    """Verifies worker returns -1 for loaded_lines when flag is active."""
    dummy_module = "sys"
    run_worker(dummy_module, skip_line_count=True)
    captured = capsys.readouterr()
    
    assert "__METRICS__:" in captured.out
    metrics_line = [l for l in captured.out.splitlines() if l.startswith("__METRICS__:")][0]
    metrics = json.loads(metrics_line.split("__METRICS__:", 1)[1])
    
    assert metrics["loaded_lines"] == -1


def test_run_worker_counts_lines_correctly(capsys):
    """Verifies worker correctly resolves paths and counts raw lines."""
    dummy_module = "math"
    mock_module = MagicMock()
    mock_module.__file__ = "/mock/path/dummy_module.py"
    dummy_content = "import os\nprint('hello')\nx = 1\n"
    
    class CustomModulesDict(dict):
        def __init__(self, *args, **kwargs):
            super().__init__(*args, **kwargs)
            self._call_count = 0
        def keys(self):
            self._call_count += 1
            if self._call_count == 1:
                return {"math"}
            return {"math", "dummy_test_mod"}
            
    fake_modules = CustomModulesDict({
        "math": sys.modules["math"],
        "dummy_test_mod": mock_module
    })
    
    with patch("sys.modules", fake_modules), \
         patch("builtins.open", mock_open(read_data=dummy_content)), \
         patch("profiler.importlib.invalidate_caches"):
         
         run_worker(dummy_module, skip_line_count=False)
         captured = capsys.readouterr()
         
         metrics_line = [l for l in captured.out.splitlines() if l.startswith("__METRICS__:")][0]
         metrics = json.loads(metrics_line.split("__METRICS__:", 1)[1])
         
         assert metrics["loaded_lines"] == 3


# =====================================================================
# 3. PARSER TESTS (_run_worker_and_parse)
# =====================================================================

def test_run_worker_and_parse_success():
    """Verifies master extracts metrics JSON block from worker stdout."""
    mock_stdout = (
        "__METRICS__:{\"time_ms\": 15.5, \"peak_ram_mb\": 12.0, \"rss_ram_mb\": 10.0, "
        "\"loaded_modules\": 22, \"loaded_lines\": 120}\n"
    )
    mock_process = MagicMock(stdout=mock_stdout, stderr="")
    
    with patch("subprocess.run", return_value=mock_process):
        data = _run_worker_and_parse(["python", "profiler.py"])
        assert data["time_ms"] == 15.5
        assert data["loaded_lines"] == 120


def test_run_worker_and_parse_forwards_stderr(capsys):
    """Verifies worker stderr warnings are forwarded to master's stderr."""
    mock_stdout = (
        "__METRICS__:{\"time_ms\": 10.0, \"peak_ram_mb\": 12.0, \"rss_ram_mb\": 10.0, "
        "\"loaded_modules\": 22, \"loaded_lines\": 120}"
    )
    mock_stderr = "DeprecationWarning: some_pkg is deprecated"
    mock_process = MagicMock(stdout=mock_stdout, stderr=mock_stderr)
    
    with patch("subprocess.run", return_value=mock_process):
        _run_worker_and_parse(["python", "profiler.py"])
        captured = capsys.readouterr()
        assert "DeprecationWarning: some_pkg is deprecated" in captured.err


# =====================================================================
# 4. MASTER COORDINATION & BENCHMARK INTEGRATIONS
# =====================================================================

@patch("profiler._run_worker_and_parse")
def test_run_master_skips_line_count_and_restores_metrics(mock_parse):
    """Verifies master appends --skip-line-count and restores metrics."""
    run_data_1 = {"loaded_modules": 5, "loaded_lines": 1000, "time_ms": 50.0, "peak_ram_mb": 10.0, "rss_ram_mb": 8.0}
    run_data_2 = {"loaded_modules": 5, "loaded_lines": -1, "time_ms": 40.0, "peak_ram_mb": 10.0, "rss_ram_mb": 8.0}
    mock_parse.side_effect = [run_data_1, run_data_2]

    with patch("builtins.print"), patch("sys.stderr"):
        run_master(iterations=2, target_module="dummy_mod", cpu=NO_CPU_PINNING, clear_cache=False)
        
        second_cmd_args = mock_parse.call_args_list[1][0][0]
        assert "--skip-line-count" in second_cmd_args
        assert run_data_2["loaded_lines"] == 1000


@patch("profiler._run_worker_and_parse")
def test_run_master_checks_non_deterministic_behavior(mock_parse, capsys):
    """Verifies warnings are printed upon non-deterministic module loads."""
    mock_parse.side_effect = [
        {"loaded_modules": 50, "loaded_lines": 500, "time_ms": 10.0, "peak_ram_mb": 1.0, "rss_ram_mb": 1.0},
        {"loaded_modules": 55, "loaded_lines": -1, "time_ms": 10.0, "peak_ram_mb": 1.0, "rss_ram_mb": 1.0}
    ]

    run_master(iterations=2, target_module="dummy_mod", cpu=NO_CPU_PINNING, clear_cache=False)
    captured = capsys.readouterr()
    assert "WARNING: Non-deterministic import behavior!" in captured.err


@patch("profiler._run_worker_and_parse")
def test_run_master_with_cpu_pinning(mock_parse):
    """Verifies taskset command configuration on Linux platforms."""
    mock_parse.return_value = {
        "loaded_modules": 10, "loaded_lines": 500, "time_ms": 12.0, "peak_ram_mb": 1.0, "rss_ram_mb": 1.0
    }
    
    with patch("sys.platform", "linux"), patch("builtins.print"), patch("sys.stderr"):
        run_master(iterations=1, target_module="dummy_mod", cpu=1, clear_cache=False)
        
        args = mock_parse.call_args[0][0]
        assert "taskset" in args
        assert "1" in args


@patch("profiler._run_worker_and_parse")
def test_run_master_writes_csv_and_diffs_baseline(mock_parse, tmp_path):
    """Verifies benchmark results CSV writing, reading, and thresholds comparison."""
    mock_parse.return_value = {
        "loaded_modules": 10, "loaded_lines": 500, "time_ms": 12.0, "peak_ram_mb": 1.0, "rss_ram_mb": 1.0
    }
    
    csv_file = tmp_path / "results.csv"
    baseline_file = tmp_path / "baseline.csv"
    
    # Write mock baseline CSV data
    with open(baseline_file, "w", newline="") as f:
        writer = csv.writer(f)
        writer.writerow(["iteration", "time_ms", "peak_ram_mb", "rss_ram_mb", "loaded_modules", "loaded_lines"])
        writer.writerow([1, 10.0, 1.0, 1.0, 10, 500])

    with patch("builtins.print"):
        run_master(
            iterations=1,
            target_module="dummy_mod",
            cpu=NO_CPU_PINNING,
            csv_path=str(csv_file),
            diff_baseline=str(baseline_file),
            diff_threshold=50.0,
            clear_cache=False
        )

    assert csv_file.exists()

Comment thread packages/google-auth/google/auth/__init__.py Outdated
Comment thread packages/google-cloud-compute/google/cloud/compute/__init__.py Outdated
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf: import-profiler redundantly reads and counts module lines on every iteration

3 participants