for Pharma
Toggle menu

FAQ

Everything a launch team asks before they hire us.

What we build, who shows up, what it costs, what happens when the date moves, and who owns the platform at the end. If your question is not here, ask it on the contact form and we will answer it directly.

The work

What we do, and what we leave to you

The short version: we hold the engineering lane, your analysts hold the business lane.

What does DataKitchen do for commercial pharma?

We engineer the commercial data platform behind a pharma product launch. That means integrating the vendor feeds, building star schemas your analysts can actually query, engineering the context layer your AI tools need, and running tests on every table. Our engineers sit in your standup through launch day and the first year. The platform runs in your cloud and transfers to your team when you say.

What do you not do?

We don't do incentive compensation, territory alignment, call planning, or brand strategy. We build the commercial data layer underneath all of them. Analytics, dashboards, and business decisions stay with your team, because that is where they belong. Staying in one lane is why a small team can move fast on it.

Who do you work with?

Two situations. Small and mid-size pharma with a launch 12 to 24 months out, where the shape is one or two assets and a commercial data lead who has not launched a brand before. And companies with a product already selling that need the data run-rate down to fund the next one. Three of our customers were later acquired by top-five pharma for $100 billion combined.

How does an engagement actually run?

Three phases. Build runs from roughly 18 months out to launch minus one, and ends with a production platform six months before launch day. Run covers launch day through year one, shipping new datasets in days and fixing breakage before the field sees it. Transfer happens whenever you say, to whoever you name.

How soon do we see something working?

First star schema in four weeks. First AI context layer in six. Full launch report set six months before go-live. Field reporting has been live six months ahead of launch day on every engagement we have run. You are not waiting two quarters for a design document.

How far before launch should we start?

For a launch, 12 to 24 months out is cheapest and calmest: the first 6 to 12 months set the drug's lifetime revenue trajectory and you get one chance at the field's trust. If a product is already selling, timing works differently. We take the existing operation over feed by feed, so there is no window you can miss.

How do you keep us in the loop?

Daily scrum and ticket triage, all tracked in Jira. A weekly status meeting with your analytics lead covering work completed, in progress, and planned. A quarterly business review with leadership covering timelines and metrics on bugs, tests, and datasets. Documentation lives in a shared wiki. Time is tracked and reported.

The team

Who shows up, and what they have done before

How many engineers do you put on a launch?

One to three, typically. Roughly 1.5 engineers built the entire Cobenfy launch data platform at Karuna over two years, work most pharma companies staff with 10 engineers. At Celgene, seven of our engineers supported a $10 billion-a-year commercial portfolio. Small senior teams backed by automation, not a bench of generalists.

Who will we actually work with?

Our founders, not an account executive. Chris Bergh, our CEO, takes the first call on a new launch himself. Eric Estabrooks is the senior engineer behind most of our pharma launches and the person you are likely to meet on a Tuesday status call. Gil Benghiat owns the practice, and Chip Bloche sets the engineering bar on every platform we ship.

Do your engineers know pharma data on day one?

Yes. Our engineers know NPP, payer claims, specialty pharmacy, syndicated data, and Veeva before they start. A generalist IT team can learn those nuances eventually, but the learning happens on your launch timeline. Nobody spends the first three months finding out what a specialty pharmacy feed is.

How is this different from ZS, Beghou, or Accenture?

Analytics consultants sell analyst hours, billed at around $300 each, and hand you decks. Accenture and Deloitte bring a 30-person engineering team and a statement of work. We send one to three engineers who write code in your repository. The endgame is different too: they keep the work, we hand it over.

How long has DataKitchen done this?

Since 2013, and commercial pharma launches since the Celgene days. The company was founded in Cambridge, MA, is bootstrapped, and is fully remote. The founders co-authored the DataOps Cookbook and the DataOps Manifesto, so the methods behind the work are published rather than proprietary.

The data

The feeds, and how they break

What data sources do you work with?

IQVIA, ICON/Symphony, Veeva, specialty pharmacy dispense feeds, medical claims, non-personal promotion, real-world evidence, copay, and EDI 852/867. There is no learning curve on any of them. We work with the data you already receive rather than asking you to standardize it first.

What actually goes wrong with commercial pharma data?

Partial files that look complete, vendor restatements that read as growth, and territory realignments that make a rep's numbers drop 40 percent overnight. None of those trip a null check, because the data that arrived is perfectly valid and simply incomplete, restated, or remapped. The way you find out is a phone call from the field.

Why is commercial pharma data harder than other data?

Because you did not create most of it, and the failures are relational rather than local. A prescriber or account record gets assembled from several vendor feeds arriving on their own schedules in their own formats. Identifiers like NPI, NDC, and DEA fail silently and quietly drop rows on the join, so every downstream number comes in low with nothing erroring.

Can you handle specialty pharmacy data?

Yes, and it is usually the messiest part of a launch. A specialty brand pulls dispense data from a dozen pharmacies or more, each with its own file, schedule, and format. None of them agree on what a ship date, a fill date, or a patient status means, and each de-identifies patients its own way. Stitching that into one patient journey is the work.

How do you know a vendor feed arrived intact?

Two checks, not one. Extract-to-load reconciliation counts rows and sums amounts against the vendor's control record, so partial and truncated files stop in staging. Then a snapshot diff compares each period against what the vendor sent last time, which is the check most teams skip and the one that catches a silent restatement before it reshapes a trend line.

What error rate should we expect?

Low enough that the field stops asking. A commercial insights leader at Celgene reduced errors to about one per quarter and held that for years. The Otezla launch team shipped hundreds of schema and dataset changes a week with negligible error rates, because tests run at every stage rather than at the end.

AI

What has to be true before AI answers a real question

Will this make Snowflake Cortex Analyst or Databricks Genie work?

Not on their own, and not until the data underneath carries context. When MIT's NLP group tested leading text-to-SQL models on the Spider 2.0 benchmark of real enterprise databases, accuracy fell from over 85 percent on academic benchmarks to under 20 percent. The model was not broken. The data underneath it had no context.

What is the context layer you build?

Schema and grain, written down. Business definitions for equalized prescriptions, territory alignment, and payer hierarchy. Validated example queries. Live freshness state on every IQVIA or ICON/Symphony refresh. Without it your AI tool guesses at which of 150 tables to join and answers confidently and wrongly, which is worse than not answering.

How do you use AI in your own engineering?

AI writes the SQL transforms and the data quality tests. AI watches the pipelines and triages failures. AI generates documentation. Our engineers spend their time on what AI cannot do: deciding what to build, designing the right data shape, and talking to your team. The framework behind it is published and open.

Our CEO wants AI in the launch deck. Where do we start?

With one data source, end to end. Pick the feed the brand team asks about most, engineer the context layer over it, and see what your AI tool answers when it is pointed at the right shape of data. That is a six-week exercise, not a strategy program, and it tells you whether the rest is worth funding.

Cost and terms

What it costs and what happens if the date moves

What does this cost?

A flat monthly rate, set at scoping. For comparison, six fully loaded commercial data FTEs in Boston or the Bay Area run $1.5 million to $2 million a year, plus a six to nine month recruiting cycle. One pharma replaced its existing setup and went from $1 million a year to $300,000, with better answers because trust went up.

What happens if our launch slips?

The rate scales down with you on 30 days' notice, with no penalty for an FDA delay you did not choose. That matters when roughly 3,500 staff left the FDA after the 2025 cuts and biotechs are missing meetings and pushing trials. A fixed multi-year commitment is the wrong shape for a date you do not control.

Should we just hire the team instead?

Eventually, maybe, and we will hand the platform to them when you do. Today the hiring cycle runs six to nine months before anyone ships, and specialized commercial pharma data engineers are hard to recruit and harder to keep. Starting now with a team that already knows the feeds buys back the recruiting window.

What do we get back beyond the platform?

Your analysts' week. Data janitorial work absorbs 30 to 50 percent of a commercial analyst bench: reconciling the IQVIA refresh against claims, chasing a payer name that arrived three different ways, rebuilding a territory roll-up after someone changed the alignment file. At Celgene, 10 to 12 analysts covered hundreds of datasets without missed SLAs because the engineering layer held.

How do we get started?

Two meetings. The first is 30 minutes: you describe the launch, the team, the data sources, and the timeline, and we tell you whether we are a fit. If we are not, we say so and recommend who is. The second is 60 minutes walking one of your data sources end to end on your screen.

Is there a way to try this before committing?

Yes. A paid two-week scoping engagement leaves you with a working context layer over one of your data sources, an AI-readiness scorecard, and a sized proposal. If you decide not to move forward, the work is yours to keep. That is the whole point of building in your cloud from the first day.

In market

Taking over what you already run

For a team with a product already selling, another asset behind it, and orders to bring the run-rate down.

We already have a platform. Do you replace it?

Usually we take it over rather than rebuild it. Your warehouse stays in your cloud account, your dashboards keep working, and we change what feeds them. Where something has to be rebuilt we say so during the assessment and give you the reason, rather than proposing a rewrite because rewriting is easier for us.

Can you work alongside our current vendor?

Yes, and it is the normal case. On the BMS CAR-T warehouse, 41 contributors from five companies including Beghou and Accenture worked the same codebase over seven years. That only works when tests are the contract, so a newcomer's change either passes or fails loudly instead of quietly breaking someone else's report.

Will our reports break while you take over?

They should not, because we add tests before we change anything. Coverage goes on the feeds first, so we can prove a change did not move a number. There is no cutover date either: your incumbent keeps running while we take feeds over one at a time, which means there is no single day when everything has to work.

Can we start with one product instead of everything?

Yes. One product, or even one feed, is a good place to start when the point is to prove the run-rate moves before you commit further. It also keeps the first phase inside a budget you already control rather than one you have to go and ask for.

How do we know the savings last?

Because the structure changes rather than the effort. A cost-reduction project ends and the run-rate drifts back. Generated tests do not decay the way a hand-written suite does, volume bounds learned from history do not go stale like fixed thresholds, and a two-stage architecture has no intermediate layers to maintain. Nobody has to police it.

Why cut data spend rather than something else?

Because it is the line you can move without touching the brand, the field, or the pipeline. Promotional spend and territory coverage protect the revenue that funds everything, and pipeline programs are the company's future. The plumbing under a product that is already selling is where money is freed without anything visible getting worse.

Ownership

Who owns the platform, and how you get it back

Do we own the platform and the code?

Yes, from day one. The warehouse sits in your cloud account and every line of code, test, and schema lives in your repository. The data quality software underneath is Apache 2.0, so there is no proprietary runtime to leave behind and no license that expires when we do. You are not renting access to something we hold.

What does the handover look like?

A normal day rather than an event. We transfer to whoever you name: your team, Accenture, or your offshore center. No lock-in, no exit fees. Eisai transferred one of our platforms to Accenture cleanly when they were ready. Because every line of code and every test already lived in their environment, there was nothing to extract.

Can other teams work on the platform alongside you?

Yes, and they do. On the BMS CAR-T warehouse, 41 contributors from five companies, including Beghou and Accenture, have worked the same codebase over seven years: 70 datasets, 253,485 lines of code, 594 data quality tests, still in production. That only works when the tests are the contract, so a newcomer's change either passes or fails loudly.

Where does our data actually live?

In your cloud account. The warehouse and the data quality testing both run there, and the tests execute as SQL inside your own warehouse, so your rows never cross the wire to us. The orchestration service runs in our cloud and drives your environment without holding your data. For a pharma security review, that split is usually the first thing to establish.

Does this hold up in acquisition due diligence?

It has, three times. Due diligence scrutinizes whether the commercial story is accurate, auditable, and defensible, and a reconciliation trail on every feed is what turns "trust me" into "here is the check that passed." Our customers' platforms survived every exit and stayed with the customer each time.

Track record

The launches behind the claims

Which pharma launches has DataKitchen worked on?

Karuna Therapeutics for the Cobenfy launch, Celgene for Otezla and Revlimid, and Acceleron. BMS acquired Karuna for $14 billion and Celgene for $74 billion. Merck acquired Acceleron for $11 billion. The platform stayed with the customer through every one of those transactions.

What results can you point to?

An average of 1.5 engineers built the whole Cobenfy launch platform at Karuna over two years, covering 50 integrated datasets. At Celgene, seven of our engineers supported $10 billion in US brand sales, and only 18 percent of engineering time went to data loading rather than new analyst work. One pharma cut data costs by two-thirds when we replaced its setup.

Where should I start reading?

Two posts cover the argument end to end. The $100 Billion Secret explains why pharma companies hand commercial data to a specialized outside team instead of a general IT queue. The cash-runway post works the numbers for a small pharma launching in 12 to 24 months.

Still have a question?

Thirty minutes with our team. Bring your launch window, your data sources, and the hardest question your brand team has asked you this quarter.