Add 2026-06-llmops-quickstart blog code - #91
Open
CEDipEngineering wants to merge 6 commits into
Open
Conversation
End-to-end LLMOps quickstart on Databricks (customer support ticket classifier): MLflow ChatAgent on a Foundation Model API endpoint, an evaluation gate that promotes a Champion in Unity Catalog, and a Databricks Asset Bundle that deploys schema, experiment, and jobs with batch + real-time inference. - Folder follows the YYYY-MM-[name] convention - README documents setup, structure, data (30 synthetic tickets, no PII), and licenses - LICENSE.md is an unmodified copy of the repo Databricks license - CODEOWNERS entry added for the new folder I have read the contribution guidelines. No sensitive info, no PII, no external dataset. Pending: SME code review approval and internal approval. Co-authored-by: Isaac
CEDipEngineering
requested review from
QuentinAmbard,
alanreese-dbrx,
alexott,
anupkalburgi,
kwulffert23,
matthewmoorcroft and
srinivasadmala
as code owners
June 5, 2026 19:07
Co-authored-by: Isaac
…al, AI Gateway logging Bring the quickstart current (2026) and align it with the MLOps quickstart's step-by-step, best-practices structure. Tested end-to-end across three workspaces (deploy → 4 jobs → approval → governed endpoint → batch + realtime). - Evaluation: replace the hand-rolled accuracy loop with mlflow.genai.evaluate() using a deterministic exact_match gate scorer plus the built-in Correctness LLM judge; every row is captured as an MLflow Trace. - Lifecycle: passing versions register as Challenger; new model_approval.py promotes Challenger→Champion only when run with --params approved=true, wired as a predecessor task to deployment. - Governance: after agents.deploy(), enable AI Gateway inference-table payload logging on the agent endpoint (idempotent). Guardrails/rate limits documented as an FM-endpoint pattern (not supported on custom agent endpoints). - Agent: normalize response content so reasoning models (Claude Sonnet 5, GPT-5) that return structured content blocks work, not just plain-string responses. - Default LLM endpoint → databricks-claude-sonnet-5; pin mlflow>=3.4.0 in the logged model. - Harden stale-deployment cleanup; README updated with the new flow, approval gate, governance notes, UC-privilege prerequisite, and a "before you call it done" checklist.
Per SME review, move off Model Serving / agents.deploy() to the current recommended pattern: serve the agent as a Databricks App and call the LLM through a Unity AI Gateway model service. Validated locally (agent server serves correctly) and against real model services on two workspaces; MLflow 3 GenAI eval passes at 90%. - agent_server/: FastAPI app (mlflow.genai.agent_server @invoke handler) classifying a ticket into one of five categories; calls the model service by fully-qualified name via the AI Gateway; reasoning-model-safe content handling. - LLM via UAIG model service (LLM_MODEL = <cat>.default.claude-sonnet-5; also gpt-oss-120b); governance lives in Unity Catalog. - Evaluation gates the app deploy: mlflow.genai.evaluate with an exact_match gate scorer + the out-of-the-box Correctness judge; agent-evaluate exits non-zero below threshold. Dropped UC model registration and Champion/Challenger aliases. - Bundle: app resource + experiment + data-ingestion job (writes the support_tickets UC managed table). uv.lock via --exclude-newer for proxy-mirrored versions. - Current naming throughout: Declarative Automation Bundles, UAIG, UC managed table, out-of-the-box. README adds prerequisite-knowledge, qs_catalog naming, training pointers, UC-privilege note, and a "before you call it done" checklist.
Link the free DevOps Essentials for Data Engineering (bundles/CI-CD) and Building Agentic Applications on Databricks (agents, MLflow, evaluation) courses, per SME suggestion to point readers at relevant training.
Co-authored-by: Isaac
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this is
End-to-end LLMOps quickstart on Databricks — a customer support ticket classifier that carries an LLM agent through the full lifecycle: data ingestion → agent build → evaluation → approval → deployment → inference.
Accompanies an upcoming Databricks Community blog post (JIRA TLC-1077).
Architecture (2026 pattern)
mlflow.genai.agent_server@invokehandler), not a Model Serving endpoint. Serving agents as apps is the current recommendation;agents.deploy()is reserved for special custom cases.<catalog>.default.claude-sonnet-5, alsogpt-oss-120b) through the AI Gateway. The model service is a Unity Catalog securable, so access control, rate limits, and payload logging live in UC. One env var (LLM_MODEL) switches models.mlflow.genai.evaluatewith a deterministicexact_matchgate scorer plus the out-of-the-boxCorrectnessjudge; every prediction is captured as an MLflow Trace.agent-evaluateexits non-zero below the threshold (CI gate). A human approves before the app is deployed.Contents
2026-06-llmops-quickstart/(YYYY-MM-[name]convention)agent_server/(agent, evaluation, server),app.yaml,pyproject.toml+uv.lockdatabricks.yml) with the app resource, MLflow experiment, and the data-ingestion job (dev/prod targets)README.mdwith prerequisite-knowledge, setup, governance recap, and a "before you call it done" checklistLICENSE.md(unchanged) andCODEOWNERSTesting
/invocationsreturns correct categories.mlflow.genai.evaluaterun: 90% exact-match, gate passed, traces logged.Guidelines checklist
YYYY-MM-[folder_name]Blog post link will be added to the README once published.
This pull request and its description were written by Isaac.