Hey readers —

We’re back with another company deep dive, where we look at how teams are actually building and deploying AI in healthcare.

This time, we’re covering Bunkerhill Health, the company behind Carebricks, an agentic AI platform that helps health systems turn their own clinical and operational ideas into production-ready agents.

Health systems have no shortage of ideas for improving care, reducing cost, or fixing broken workflows, but what they usually lack is the time, technical infrastructure, and operational capacity to put those ideas into practice. Bunkerhill helps them do exactly that.

In this deep dive, we cover why dashboards often let hospitals “admire the problem” instead of solving it, how Carebricks moves from insight to action, why Nish believes most healthcare AI use cases are becoming “LLMs plus X,” and what it takes to build a real platform rather than a bundle of point solutions.

Let’s dive in. 👇

Read time: 8 minutes

TOGETHER WITH BUNKERHILL

Company Deep Dive: Bunkerhill

Perspectives from the people building the future of health AI…

We sat down with Nish Khandwala, CEO and Co-Founder of Bunkerhill Health, to discuss how the company is helping health systems move AI from a promising idea into something that actually runs in production.

Bunkerhill’s platform, Carebricks, gives hospitals a shared foundation for building and deploying AI agents across clinical, operational, and administrative workflows. Those agents can review records, reason across clinical and enterprise data, communicate with patients and providers, create referrals and orders, and write back into systems like EHRs, payer portals, registries, and other tools.

The company works exclusively with health systems, including Cleveland Clinic, the University of Texas Medical Branch (UTMB), and Intermountain Health. More than 20 agents are already live at UTMB across areas like referral triage, incidental findings, registry management, and other clinical and operational workflows.

Bunkerhill recently raised a Series B led by Khosla Ventures, bringing total funding to $55M, with continued participation from Sequoia Capital, Felicis, Optum Ventures, and Y Combinator.

Nish shared how the company grew out of a failed attempt to operationalize a preventive cardiology idea at Stanford, why identifying a problem is not the same as solving it, how Carebricks combines knowledge, reasoning, and action, and why the long-term opportunity is helping health systems turn far more of their best ideas into reality.

Let’s start from the top. What sparked Bunkerhill Health into existence?
We did not start Bunkerhill because we set out to start a company. It grew out of something we lived through.

My co-founder, David, and I were computer science students at Stanford focused on applied AI in healthcare. We saw a lot of great ideas for how AI could improve care. One came from the head of preventive cardiology at Stanford: could we use scans that had already been taken for other reasons to identify patients at risk of heart disease and bring them into preventive care before they had a heart attack?

It seemed like a no-brainer. The AI itself did not feel like the hardest part. But even with the cardiologist and the resources of Stanford behind it, we spent two years trying and failing to translate that idea into clinical practice.

Around that time, my dad had a heart attack. He is fine now, thankfully, but we later learned that an earlier scan had already shown signs of blockages in his arteries and nobody had acted on it. The exact idea we had struggled to implement could potentially have prevented what happened to him.

That made the problem very real. Healthcare has no shortage of good ideas, whether they are meant to improve outcomes, lower costs, or make operations more efficient. What it struggles with is turning those ideas into something that actually runs. Bunkerhill was born out of that frustration.

For readers hearing about Bunkerhill for the first time, what does the company do today?
Bunkerhill helps health systems turn their own clinical and operational ideas into production-grade AI agents.

Carebricks is the shared foundation underneath those agents. A clinical leader, operational leader, or administrative team can bring us a problem, and the platform gives them the ability to create an agent that handles the work end-to-end.

Our only customer type is the health system. We work with more than a dozen today, including Cleveland Clinic and the University of Texas Medical Branch. At UTMB alone, 22 agents are already live in production across patient care, operations, and administration.

The goal is not to show health systems what AI could theoretically do. It is to help them take the ideas they already have and make those ideas operational.

How do you think about the different kinds of AI agents a health system can build?
AI is ultimately a tool for meeting different needs that a hospital has. The way I picture it is a hierarchy, a kind of Maslow's hierarchy for AI use cases, where the needs a hospital cannot ignore sit at the base and the more ambitious ideas sit higher up.

At the base of the pyramid are the hair-on-fire problems. These are problems the hospital cannot afford to ignore. They may already be preparing to throw people at the issue manually, but they would much rather make it go away.

That could be patients not receiving medication quickly enough. It could be a change in Medicaid definitions that puts hundreds of millions of dollars in rebates or discounts at risk. These are not nice-to-have projects. If they are not addressed, the consequences can be existential.

As you move higher up the pyramid, you get into frontier use cases. They are still ROI-positive and good for patient care, but the hospital can technically afford to wait. Doing them puts the organization ahead of the market, but not doing them tomorrow does not immediately put the hospital in crisis.

Health systems usually bring us in for something near the base because that is where the urgency and budget are. Once the platform is in place, people from across the organization start raising their hands with ideas from higher up the pyramid.

Ironically, solving the urgent foundational problems creates the path for more ambitious clinical innovation.

What do people consistently misunderstand about building AI for health systems?
The biggest misunderstanding is the assumption that every healthcare AI use case still requires a bespoke solution.

That was true when AI was narrow. A prior authorization product, registry abstraction product, and care gaps product each needed a different algorithm, technical stack, and vendor. The options for a health system were to wait for an incumbent, buy another point solution, build something custom, or keep throwing people at the problem.

Large language models change that strategy. For the first time, we have a genuinely general technology that can read plain English, understand context, and use tools. More and more use cases now look like LLMs plus X, where X is the clinical model, integration, workflow logic, or action layer needed for that specific problem. As the general models improve, X keeps getting smaller.

If you believe that, you stop thinking about AI as a long list of unrelated purchasing decisions. You start thinking about the common foundation that can support many use cases.

That mindset shift is difficult because this is the first genuinely general-purpose technology most healthcare leaders have encountered. We are used to software being narrow. It takes a different way of thinking to realize that the same underlying platform may be able to handle prior authorization, registry management, referral triage, care gaps, and use cases that have not even been proposed yet. A single foundation can deliver the breadth of an enterprise-wide platform with the depth of purpose-built point solutions.

You describe Carebricks as an end-to-end agentic platform rather than a copilot. What is the practical difference?
The difference is whether the AI identifies work or actually completes it.

Take care gap identification. An AI tool might review historical echocardiograms and identify 1,000 patients with severe aortic stenosis who never received appropriate follow-up.

A copilot gives the hospital that list. But a large health system cannot snap its fingers and find 10 or 15 nurses to contact 1,000 patients. The tool may have found the right answer, but no patient receives better care because the work still has not been done.

An end-to-end agent identifies those patients, sends MyChart messages and letters, creates referrals, supports scheduling, and moves them into the right clinical workflow. The hospital does not receive another task. The agent carries out the task.

The same applies to registry abstraction. A copilot can find the answers in the chart, but someone still has to copy and paste them into the registry portal. An end-to-end system writes the answers back into the portal.

Dashboards allow hospitals to admire the problem. Agents are supposed to solve it.

Can you share a few examples of the impact Carebricks is having today?
Cleveland Clinic is a strong example.

A few years ago, they created an actionable findings clinic to make sure patients with unexpected findings in radiology reports received appropriate follow-up. They staffed it with roughly 10 nurses, which was a major commitment to patient safety.

The challenge was volume. A backlog started building, and for a workflow involving possible cancers or other serious findings, a backlog can be almost as bad as not doing the work at all.

Our agents were designed to mirror the clinic’s existing process. Cleveland Clinic eliminated the backlog and is now able to move 44% more patients through the actionable findings clinic with the same staff. That is operational efficiency, but it is also better patient care because more people receive follow-up earlier.

At UTMB, one agent triages nephrology referrals based on severity. After deploying the agent, the average wait time for a patient at UTMB to see a nephrology specialist was reduced by over 50%.

That is the type of impact we care about. The agent is not simply generating an insight. It is helping sicker patients get care sooner.

What does the technology stack look like under the hood?
We think of Carebricks as a system of action that sits on top of systems of record.

The platform has three primary “bricks”, or building blocks. The first is the knowledge brick, which connects with the hospital’s EHR, ERP, supply chain system, imaging archive, EKG archive, external portals, and even the broader internet where appropriate.

The second is the reasoning brick. General reasoning comes from frontier language models from companies like OpenAI, Anthropic, and Google. Alongside them sits a library of proprietary clinical algorithms licensed from academic institutions, available to any agent whose use case calls for specialized clinical interpretation. A referral triage agent may never touch one. An imaging agent might depend on it.

We have partnerships with 27 academic centers, including Stanford, Cleveland Clinic, and UCSF. Over the past 12 months, we have obtained seven FDA clearances for models licensed through those relationships. It is somewhat similar to a pharma licensing model, except we are licensing clinical AI algorithms and sharing revenue with the institutions that developed them.

The third is the action brick. That is what allows the system to communicate with patients and providers, create orders and referrals, write back to the EHR, send messages or letters, make voice calls, and enter information into external systems like CoverMyMeds, payer portals, and clinical registries.

The value comes from combining knowledge, reasoning, and action into a reusable system rather than rebuilding the whole stack for every use case.

Clinical teams are cautious about handing workflows to AI. How do you earn trust and decide where humans stay involved?
Each new use case pairs the problem owner at the health system with a forward-deployed lead from Bunkerhill.

That person is responsible not only for building and monitoring the agent, but also for the change management around the deployment. We do not move from an offline prototype to enterprise-wide automation overnight.

We first test performance against the standard of care using historical data. Then we might launch with one user or a subset of patients. Once the health system sees that the agent performs well against the status quo, we expand gradually.

We also require traceability. When the AI produces an output, it needs to cite the information in the system of record that supports it. It should not be possible for an agent to make a claim without showing where that claim came from.

From there, we monitor downstream outcomes and use additional AI systems to evaluate how the agents are performing. The rollout is iterative and done in lockstep with the clinical or operational team that owns the problem.

A lot of digital health companies call themselves platforms. What makes Carebricks a true platform rather than a collection of point solutions?
This is a pet peeve of mine. Most companies that call themselves platforms are really multi-point solutions.

A point solution solves one problem. A multi-point solution solves a handful of predefined problems. A true platform supports use cases that were not known when the product was originally built.

If you asked me how many use cases Carebricks can support, I could not give you a fixed number. New use cases are being proposed and built all the time. That is closer to the iPhone App Store, where the range of possible applications is not predetermined.

The key is abstraction. Most agents can be assembled from some combination of a knowledge brick, reasoning brick, and action brick. Once those abstractions exist, teams can connect the pieces differently to create new workflows.

The test is simple: the 100th use case should be much easier to build and deploy than the first. If every new project requires starting over, you do not have a platform. You have a consulting business.

What metrics matter most when you evaluate whether Carebricks is working?
At the health system partner level, the main input metric is the number of use cases that are actually live in production.

We do not index heavily on the number of users because many agents do not have a traditional user. We also do not care about maximizing token usage. Burning more tokens does not mean more value is being created.

The output metric is ROI. How much money did the agent make the hospital, and how much did it save?

On the revenue side, that can mean identifying more patients who need appropriate downstream procedures or helping the health system capture reimbursement it was previously missing.

On the cost side, we look at vendor consolidation, reductions in third-party professional services contracts, and whether the organization can handle more work with the same number of people.

The point is not AI usage for its own sake. It is whether the health system can point to a measurable improvement in care, capacity, revenue, or cost.

Five years from now, what changes if Bunkerhill executes the way you hope?
I think about that from three perspectives: the person with an idea, the health system, and the patient.

For the person with an idea, the expectation should change completely. I want people to complain that they had an idea yesterday and it still is not live across the enterprise today. The new normal should be going from an idea to production in less than 24 hours.

From the health system’s perspective, I would like to see hospitals running on Carebricks across clinical operations, administrative work, and revenue cycle. Humans should spend their time on the highest-leverage work, not on repetitive tasks that AI can reliably execute.

For patients, the result should be faster access to the best available care. Healthcare should no longer feel like an industry that is always a decade behind, still mailing CDs and relying on fax machines.

The goal is for healthcare to move from being a laggard in technology adoption to being one of the first industries to put new innovations into practice. Patients should be able to trust that when better evidence, treatments, or technology become available, their health system can act on them quickly.

Healthcare AI Guy Summary

What stood out, what’s tricky, and why it matters…

Bunkerhill Health is going after a problem that sits underneath a lot of healthcare AI: health systems usually do not lack ideas. They lack the infrastructure and operating capacity to put those ideas into practice.

Hospitals already know where patients fall through the cracks, where staff spend hours on repetitive work, and where better follow-up could improve outcomes. The harder part is connecting the data, applying the right reasoning, carrying out the next steps, and doing all of that safely inside existing systems.

Carebricks is Bunkerhill’s attempt to turn that process into a platform. Its agents pull information from systems of record, reason across it using frontier models and licensed clinical algorithms, then take action by messaging patients, creating referrals, writing back to the EHR, or completing work inside external portals. The pitch is less about giving teams another answer and more about using agents to actually finish the job.

Bunkerhill did not begin with the full platform vision. The company initially built infrastructure to move individual clinical AI models from research into practice, including sourcing data, validating performance, supporting FDA clearance, and deploying them inside hospitals. That experience exposed the next bottleneck: health systems did not have the time or resources to repeat that process for one model and one vendor at a time. Carebricks grew out of the need for a common foundation that could support a whole host of ideas and workflows.

The company has meaningful traction behind that thesis. As outlined, Bunkerhill works with more than a dozen health systems, including Cleveland Clinic, UTMB, and Intermountain Health. UTMB has 22 agents live across clinical, operational, and administrative workflows, expanding from one specialty into others as teams saw what the platform could do. Bunkerhill has also built partnerships with 27 academic centers and secured seven FDA clearances over the past year for licensed clinical models. Now, to build on the momentum, the company raised a Series B, led by Khosla Ventures, bringing total funding to $55M.

What stands out is the breadth of what health systems are already using it for. At Cleveland Clinic, Carebricks helped eliminate a backlog in the actionable findings clinic and allowed the same team to move 44% more patients through the workflow. At UTMB, an agent reduced nephrology wait times by more than 50%. Those are useful examples because the value is not measured in model outputs or token usage. It shows up in real patients receiving follow-up and care sooner.

What stood out

  • Action over insight: Nish’s line that dashboards let hospitals “admire the problem” gets to the heart of the product. Identifying 1,000 patients who need follow-up is not useful if the hospital still needs to find staff to contact and schedule all of them. Carebricks is designed to carry the work through vs. add another queue.

  • A real platform test: Bunkerhill draws a clear line between platforms and multi-point solutions. The key question is whether the 100th use case becomes easier to build than the first. The knowledge, reasoning, and action bricks provide a shared foundation that can be reconfigured instead of rebuilt for every workflow.

  • The health system supplies the ideas: Bunkerhill is not arriving with a fixed list of AI products and trying to find buyers for them. The use cases come from clinical and operational leaders inside the health system. That gives the platform a path from urgent, hair-on-fire problems into more ambitious ideas that previously lacked the staff or budget to pursue.

  • “LLMs plus X” changes the buying strategy: Prior auth, registry abstraction, care gaps, and referral triage once required separate technical stacks. As more of those tools become a general model plus a smaller amount of clinical logic, integration, and workflow configuration, a common platform starts to make more sense than another wave of disconnected vendors.

  • Abundance is the more interesting outcome: Efficiency is part of the story, but the larger opportunity is enabling work that was not happening at all. Finding incidental disease, applying new treatment guidelines across an entire patient population, or following every actionable radiology finding are meaningful because they expand what a health system can actually do.

What’s tricky

  • Breadth can turn into services work: Supporting an open-ended set of use cases is a strong platform story, but each health system has different data, workflows, governance, and internal politics. Bunkerhill will need to keep proving that new agents become faster and more repeatable to deploy rather than requiring a lot of custom configuration each time.

  • Taking action raises the stakes: Writing a summary or surfacing a list is one thing. Messaging patients, prioritizing referrals, and creating orders can directly affect care. The company’s phased rollouts, source citations, monitoring, and forward-deployed model are sensible, but expectations will keep rising as agents become more autonomous.

  • The platform has to stay as deep as the point solutions: The appeal of consolidation is clear, but health systems will not accept weaker performance just to reduce vendor count. Bunkerhill’s challenge is delivering broad infrastructure while matching the clinical and operational depth of specialized products in each use case.

Final thoughts

Bunkerhill believes the biggest challenge in healthcare AI isn’t a shortage of models, dashboards, or projects. It is the divide between recognizing what should happen and having the capacity to make it happen. Carebricks is built around bridging these two. The EHR remains the system of record, while Bunkerhill tries to become the system of action sitting across clinical, administrative, and operational work.

The broader goal is a health system where a strong idea can move from one clinician's head into production quickly, and where the same foundation can support hundreds of agents without adding hundreds of vendors. Early deployments suggest that approach is working.

The test will be whether Bunkerhill can preserve that depth and reliability as the number of agents, health systems, and clinical actions grows. If it can, Carebricks could become the infrastructure that lets health systems move from experimenting with AI to actually natively deploying and operating with it.

This issue is presented in partnership with the featured company.

That’s it for this deep dive friends! Back to reading — I’ll see you next Tuesday.

Stay classy,

— Healthcare AI Guy (X/Twitter | LinkedIn)

PS. I write this newsletter for you. So if you have any suggestions or questions, feel free to reply to this email and let me know

How was this week's newsletter? Tap your choice below👇

Login or Subscribe to participate


You may also like these