Every few months, Databricks ships another headline feature for Genie Code. At the Data + AI Summit 2026 it was a full-page command center for multi-threaded ML work, an ontology that learns how your team builds features, and autonomous overnight runs that check pipelines and summarize results while everyone sleeps. But the number that stuck with us came from the original Genie Code launch earlier in 2026: on real-world data science tasks, Genie Code’s success rate jumped from roughly 32% to 77% compared to leading coding agents. The pitch is that AI is moving from assisting data teams to actually doing the work, with humans supervising rather than typing.
Most organizations are still watching that shift happen from the sidelines, waiting for the tooling to mature before committing. We didn’t wait. Over the past several months, we’ve been running exactly this model in production for a client, a global leader in HVAC systems manufacturing, and the results are making us rethink what AI-assisted data engineering should actually mean.
The problem: rich data, real complexity, real stakes
This client’s data doesn’t live in one clean place. It’s spread across Oracle ERP, Salesforce, several other CRM applications, and SharePoint, feeding analytics programs across supply chain, quality, aftermarket, and equipment sales and service. Each of those domains has its own business logic, its own quirks in the source systems, and its own definition of what good data looks like. Multiply that by the number of gold-layer data models needed to serve each business area, and you get a delivery pipeline where the bottleneck isn’t compute. It’s the human cycles spent understanding the data, mapping it correctly, building it right the first time, and keeping it healthy once it’s live.
We had already solved the ingestion and storage side of this with a metadata-driven architecture: Azure Data Factory pulling from every source system into a medallion architecture (bronze, silver, gold) on Databricks. What we hadn’t solved was the layer above the pipelines: the analysis, design, development, review, and support work that still ran at human speed.
The approach: Genie Code as a teammate across the full lifecycle, not a single step
Most teams that adopt an AI coding agent point it at one stage, usually writing code faster. We took a different approach and embedded Genie Code across every stage of delivery, each with a distinct job:
Business analysis
When a new use case comes in from the business, Genie Code helps our analysts find the relevant tables in a genuinely large, sprawling database, apply domain context to what those tables actually mean, and produce source-to-target mapping documents. It also recommends the gold data model design itself, turning what used to be days of manual data discovery into a guided first draft.
Development
Once the design is set, Genie Code assists developers with code generation and, just as importantly, with end-to-end unit testing, automatically generating a results report rather than leaving that verification step to be squeezed in at the end of a sprint.
Code review
This is where it gets more interesting than a typical AI coding assistant. We’ve built custom skills for Genie Code that encode our own design and development standards, so every suggestion, from table naming to transformation logic, reflects how we build, not a generic best practice. Those same custom skills are then applied during code review, where Genie Code checks for issues that are easy for a human reviewer to miss after the tenth pull request of the day.
Production support
After deployment, Genie Code helps troubleshoot issues directly against production data and produces a complete root cause analysis report, turning what used to be a multi-hour investigation into a documented starting point for the engineer on call.
At every one of these stages, a human makes the final call. Genie Code is the second pair of eyes and the first draft, never the last word.
Why this is more than “AI writes code faster”
The interesting part isn’t that an AI agent can generate a pipeline or fix a bug. Plenty of tools do that now. It’s that we’ve turned Genie Code into a consistent institutional memory that spans the entire delivery chain: it carries business domain context into analysis, carries design standards into development, carries those same standards into review, and carries production knowledge into support. New engineers ramping onto this codebase lean on it to understand existing code and database structure far faster than reading documentation or shadowing a senior teammate ever allowed.
This is also how we think about our work in general. We don’t plug in a tool and bill the hours. We build the layer that makes the tool ours: the skills, the standards, and the guardrails that turn a generic agent into something that works the way we work. It is the same principle behind Aetherion, our own AI platform: let the platform do the heavy lifting, then wrap it in the engineering judgment that makes the outcome production-grade. The engineering judgment we build around it is what makes the outcome production-grade, and what lets us own that outcome instead of handing over a pile of scripts.
That’s also exactly the direction Databricks itself is pointing with the Genie Ontology and command center updates announced this year: an agent that learns how a specific team works, not just how to write generic code. We got there early by building the custom skills layer ourselves, tailored to one client’s standards rather than waiting for a generic version.
The business outcome
The result has been a measurable shift in delivery velocity and team resilience, not just a productivity anecdote:
Simple, well-understood use cases now see roughly 50 to 60% faster delivery, while complex, multi-domain builds still see a solid 30 to 40% improvement. Work that used to take a sprint now clears in less time, largely because the analysis and testing bottlenecks, not the coding itself, were where time actually disappeared.
Quality has gone up alongside speed, not at its expense, because the code review and testing stages are now checked consistently, every time, against the same standards.
Team ramp-up has gotten dramatically faster. New members can get context on existing code and database structures on their own, cutting the dependency on tribal knowledge that usually slows down onboarding in complex, multi-source data environments.
What this means if you’re evaluating agentic AI for your data org
The mistake we see most engineering leaders making is treating tools like Genie Code as a coding accelerator to bolt onto the development stage. The bigger unlock is treating it as a lifecycle participant, one that needs the same investment in training (via custom skills and standards) that you’d give a new senior hire, and the same human oversight you’d expect for any teammate handling production systems.
The technology clearly isn’t finished maturing. But the organizations that figure out how to operationalize it, skills, standards, human checkpoints and all, across analysis, development, review, and support are going to compound that advantage long before the tooling itself catches up.
If you’re a data or engineering leader wrestling with fragmented sources, growing gold-layer sprawl, or onboarding time that never seems to shrink, this is exactly the kind of problem we like to take on. Let’s talk about where AI agents actually belong in your delivery lifecycle, one stage or all of them.
Reach out to the Calfus Data and Analytics team to understand the magic we can work with your data. Where to find it, what to do with it and the rich business outcomes that can be unlocked with this gold mine.