Skip to content

Add 2026-06-llmops-quickstart blog code - #91

Open
CEDipEngineering wants to merge 6 commits into
databricks-solutions:mainfrom
CEDipEngineering:add-2026-06-llmops-quickstart
Open

Add 2026-06-llmops-quickstart blog code#91
CEDipEngineering wants to merge 6 commits into
databricks-solutions:mainfrom
CEDipEngineering:add-2026-06-llmops-quickstart

Conversation

@CEDipEngineering

@CEDipEngineering CEDipEngineering commented Jun 5, 2026

Copy link
Copy Markdown

What this is

End-to-end LLMOps quickstart on Databricks — a customer support ticket classifier that carries an LLM agent through the full lifecycle: data ingestion → agent build → evaluation → approval → deployment → inference.

Accompanies an upcoming Databricks Community blog post (JIRA TLC-1077).

Architecture (2026 pattern)

  • Agent served as a Databricks App — a FastAPI agent server (mlflow.genai.agent_server @invoke handler), not a Model Serving endpoint. Serving agents as apps is the current recommendation; agents.deploy() is reserved for special custom cases.
  • LLM via a Unity AI Gateway model service — the agent calls the model service by its fully-qualified name (<catalog>.default.claude-sonnet-5, also gpt-oss-120b) through the AI Gateway. The model service is a Unity Catalog securable, so access control, rate limits, and payload logging live in UC. One env var (LLM_MODEL) switches models.
  • MLflow 3 GenAI evaluation gates the deploymlflow.genai.evaluate with a deterministic exact_match gate scorer plus the out-of-the-box Correctness judge; every prediction is captured as an MLflow Trace. agent-evaluate exits non-zero below the threshold (CI gate). A human approves before the app is deployed.

Contents

  • New folder 2026-06-llmops-quickstart/ (YYYY-MM-[name] convention)
  • agent_server/ (agent, evaluation, server), app.yaml, pyproject.toml + uv.lock
  • A Declarative Automation Bundle (databricks.yml) with the app resource, MLflow experiment, and the data-ingestion job (dev/prod targets)
  • README.md with prerequisite-knowledge, setup, governance recap, and a "before you call it done" checklist
  • LICENSE.md (unchanged) and CODEOWNERS

Testing

  • Agent server built and run locally; /invocations returns correct categories.
  • Model services confirmed on two workspaces (AWS + Azure) through the AI Gateway.
  • mlflow.genai.evaluate run: 90% exact-match, gate passed, traces logged.
  • Bundle validates and deploys (app resource + experiment + ingestion job).
  • Note: the Databricks Apps build installs from the workspace package proxy, which can return transient package-download errors; retry the deploy if a build fails on one.

Guidelines checklist

  • I have read the contribution guidelines
  • No sensitive information / no secrets
  • No PII; no external dataset (30 small, hand-written synthetic tickets)
  • Licenses listed in the README; Databricks license unchanged
  • Folder renamed to YYYY-MM-[folder_name]
  • Added myself to CODEOWNERS
  • Domain SME has reviewed the code and approved in a PR comment (pending)

Blog post link will be added to the README once published.

This pull request and its description were written by Isaac.

End-to-end LLMOps quickstart on Databricks (customer support ticket
classifier): MLflow ChatAgent on a Foundation Model API endpoint, an
evaluation gate that promotes a Champion in Unity Catalog, and a
Databricks Asset Bundle that deploys schema, experiment, and jobs with
batch + real-time inference.

- Folder follows the YYYY-MM-[name] convention
- README documents setup, structure, data (30 synthetic tickets, no PII), and licenses
- LICENSE.md is an unmodified copy of the repo Databricks license
- CODEOWNERS entry added for the new folder

I have read the contribution guidelines. No sensitive info, no PII, no
external dataset. Pending: SME code review approval and internal approval.

Co-authored-by: Isaac
CEDipEngineering and others added 5 commits June 5, 2026 16:13
…al, AI Gateway logging

Bring the quickstart current (2026) and align it with the MLOps quickstart's
step-by-step, best-practices structure. Tested end-to-end across three workspaces
(deploy → 4 jobs → approval → governed endpoint → batch + realtime).

- Evaluation: replace the hand-rolled accuracy loop with mlflow.genai.evaluate()
  using a deterministic exact_match gate scorer plus the built-in Correctness LLM
  judge; every row is captured as an MLflow Trace.
- Lifecycle: passing versions register as Challenger; new model_approval.py promotes
  Challenger→Champion only when run with --params approved=true, wired as a
  predecessor task to deployment.
- Governance: after agents.deploy(), enable AI Gateway inference-table payload logging
  on the agent endpoint (idempotent). Guardrails/rate limits documented as an
  FM-endpoint pattern (not supported on custom agent endpoints).
- Agent: normalize response content so reasoning models (Claude Sonnet 5, GPT-5) that
  return structured content blocks work, not just plain-string responses.
- Default LLM endpoint → databricks-claude-sonnet-5; pin mlflow>=3.4.0 in the logged model.
- Harden stale-deployment cleanup; README updated with the new flow, approval gate,
  governance notes, UC-privilege prerequisite, and a "before you call it done" checklist.
Per SME review, move off Model Serving / agents.deploy() to the current recommended
pattern: serve the agent as a Databricks App and call the LLM through a Unity AI
Gateway model service. Validated locally (agent server serves correctly) and against
real model services on two workspaces; MLflow 3 GenAI eval passes at 90%.

- agent_server/: FastAPI app (mlflow.genai.agent_server @invoke handler) classifying a
  ticket into one of five categories; calls the model service by fully-qualified name
  via the AI Gateway; reasoning-model-safe content handling.
- LLM via UAIG model service (LLM_MODEL = <cat>.default.claude-sonnet-5; also
  gpt-oss-120b); governance lives in Unity Catalog.
- Evaluation gates the app deploy: mlflow.genai.evaluate with an exact_match gate
  scorer + the out-of-the-box Correctness judge; agent-evaluate exits non-zero below
  threshold. Dropped UC model registration and Champion/Challenger aliases.
- Bundle: app resource + experiment + data-ingestion job (writes the support_tickets
  UC managed table). uv.lock via --exclude-newer for proxy-mirrored versions.
- Current naming throughout: Declarative Automation Bundles, UAIG, UC managed table,
  out-of-the-box. README adds prerequisite-knowledge, qs_catalog naming, training
  pointers, UC-privilege note, and a "before you call it done" checklist.
Link the free DevOps Essentials for Data Engineering (bundles/CI-CD) and Building
Agentic Applications on Databricks (agents, MLflow, evaluation) courses, per SME
suggestion to point readers at relevant training.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant