S1E7 Scaling AI (ft. Paul Testa, NYU Langone)

 

 

 

 

Healthy Uptime Podcast Paul Testa (NYU Langone) _ Jordan Cooper (Rackspace)-20260930_150305-Meeting Recording
September 30, 2026, 7:03PM
22m 36s

Jordan Cooper started transcription

Jordan Cooper   0:03
We're here today with Dr. Paul Testa from NYU Angone, the Chief Health Informatics Officer and an emergency medicine physician. Paul, thank you so much for joining us today.

Testa, Paul   0:12
Thanks so much.

Jordan Cooper   0:13
As background for our listeners, NYU Langone Health System is headquartered in New York, New York with 2,000 beds across 7 hospitals supported by 6,000 physicians. We're going to have a lot of fun today, have a lot of different topics to cover, but let's start out with Gen. AI because why not? It's 2026. So you rolled out one of the nation's first privately managed, secure and HIPAA compliant GPT-4 ecosystems in a healthcare organization, and you did so enterprise
wide on Azure. How do you select that model because everyone across the country is trying to figure it out, that hyperscaler in particular, and what changed between the pilot and the full rollout?

Testa, Paul   0:49
Great question, particularly the last little bit.
Folks a lot smarter than me made the platform decision, but Microsoft has been an incredibly nimble partner with us. Let's be honest, we're committed to all the foundation models and all the hosts. So we have presence with all three and all three. What has been the real, the kind of the privilege in my role, and particularly at NYU Langone, is
We never had to send out the no don't e-mail, right? 36 months ago, everything gets released. We see everyone using it. We see the traffic. And the e-mail that gets sent out from leadership about generative AI is, yes, this is an incredibly powerful tool. It's going to change a lot of things. If you want to do cool things, let us know.
And we were able to partner really quickly with OpenAI and Microsoft in getting out a secure platform to say, and here's your playground. Do whatever you want here. It's HIPAA compliant. It will hold our privileged information and protected. And then if you discover cool things,
Come and talk to us and we'll help you scale. Because innovation without scale is silly. It's A distraction. It actually only matters if you can scale it.

Jordan Cooper   2:03
So I think what a lot of systems have found is they roll something out. I mean, the most popular and easy one to discuss is ambient AI, but you know, that perhaps is the most covered and therefore least interesting. But as you roll out different kinds of AI platforms across healthcare delivery systems,
I think a lot of people in your position find that there is high demand and many use cases, many users, but you have to balance that with the high tokenization cost. So how do you choose which use cases to prioritize and which users to grant access to? And how do you kind of limit their access so that you don't break the bank?

Testa, Paul   2:42
Yeah. Also, right question. And then there's that piece that came out from Bain recently on looking at the tokenization of OpEx and CapEx, right? That's going to change our calculus. We went in, Ambient maybe isn't, the conversation's not finished yet because of nursing.

Jordan Cooper   2:59
Mhm.
Mhm.

Testa, Paul   3:01
Right. Yeah. I mean, and for providers, we get it. It works incredibly well. Most patients are very accepting of it. Providers say, don't, you know, it has changed the way we work. But nursing in the inpatient setting is a very different literal conversation. They're not used to speaking out loud. And there is, I won't say resistance, but there's

Jordan Cooper   3:17
Mhm.

Testa, Paul   3:20
caution as we get our nurses engaged. So we haven't had that huge uptick yet and we're learning more what that means for nursing. I actually think it's going to be a bigger game changer for nurses than it is for, you know, LIPs and APPs. That said, we've been, we've had some lessons learned about

Jordan Cooper   3:22
Yeah.

Testa, Paul   3:41
about, you know, making sure we've got triggers when we find things going wild. We gave Ultraviolet AI, our secure multi-model portal, to 52,000 members of the workforce. We have not woken up with, you know, essentially everybody has about 5 bucks in tokens a month. Most people...

Jordan Cooper   3:46
Mhm.

Testa, Paul   4:01
Many of that goes unused and is provisioned to others. But what we have been doing is grabbing, we want transparency. So we grab the top, you know, 10 users and look at, first off, we have a limit set, so we won't, we'll stop them before they blow through anything and start generating millions or hundreds of thousands of dollars worth of.
bills and we engage with them. Like what are you trying to do and why and how do we do this differently to make sure it's scalable? And hey, this is a great idea, we'll pay for it.

Jordan Cooper   4:31
So as the AI usage grows, are you thinking about cost predictability at all? I guess, and the reason I ask is there are ways, there are some ways when you're managing CapEx, it could be that you just have unlimited
tokens or you have a limit and you say you can hit, you know, this many tokens over this amount of time, but there are other ways to kind of...
Spread the cost and have like a...

Testa, Paul   4:58
Well, there's a stage before that, right? There's a moment before that to say, you're scaling really quickly, user number 300, and it seems like you're doing incredibly cool stuff. Can we talk and understand? And maybe we've got a cheaper way for you to do this or do it on a different platform. But.

Jordan Cooper   5:00
Yeah.
Mhm.
So do you have any examples of cool stuff that you've come across?

Testa, Paul   5:17
I think.
Well, I think...
One thing that's been incredibly powerful for us is we did enter an agreement with NVIDIA in 2017, give or take. So we have a high performance computing cluster that is essentially the largest high performance computing cluster dedicated to life sciences in the world. And that cluster allows us some of our own, we can absorb costs in that.
So chip access has been incredibly important. We continue to do large investments there. When you say interesting, do you mean like use cases? Because we got tons of those.

Jordan Cooper   5:53
Yeah, I mean, well, it's interesting from two ways. Either one, interesting clinical or business use cases, or two, interesting ways to absorb kind of the GPU costs. And I think that's fascinating. You beat before the stock exploded with Gen. AI, you have access to your own Nvidia, I guess, GPU chips. It could be CPU, but

Testa, Paul   5:54
Yeah.
Yep, yep, yeah, we do, but you know everybody can't do that, and that that's that ship has sailed, so where can we use CPU? Absolutely, but another advantage I think we have is we've been able to develop our own large language model with Eric Orman's work again in Ultraviolet, which is the...

Jordan Cooper   6:16
Yeah.
Mhm.

Testa, Paul   6:33
a massive model. I think China and the Emirates have a larger model based purely on clinical documentation. So that, that's our own model hosted domestically, makes iteration much cheaper. And there are a lot of open source models, like we all have to source healthcare.

Jordan Cooper   6:47
Mhm.

Testa, Paul   6:52
We are nonprofit entities, so we're gonna have to start looking at what can we and where can we look to other models, besides the big and frankly somewhat expensive, you know, frontier models. That said, we're doing a bunch of things that are turning out to be a whole lot cheaper than we thought they would be. The thing that excites me, one of the things that excite me the most.

Jordan Cooper   6:53
Mm.

Testa, Paul   7:11
is on, we have a proof of concept right now where we are generating a summary of what happened to you in the hospital that day.
So many transactions, the velocity of transactions that occur for a particular patient in the ICU, for example, they cannot keep track of what happened to them. They cannot keep track of who they saw, what medications they had, what's going to happen to them tomorrow, what happened to them yesterday. So we have this extracted model that is looking at all this.

Jordan Cooper   7:24
Mhm.

Testa, Paul   7:39
the results that come in, the documentation at 4:00 today and says, here's the summary of what happened to you today. Right now we have a clinician reading each one and signing off on it, and then releasing it into our NYU Lingo Health app, which goes to the patient's proxies. So they've got a summary, hey, today, this is what happened to you. You got questions, let's talk about that.

Jordan Cooper   7:48
Mm-hmm.
Dialing into that for a second, and we have a lot of topics to cover, but this is just really interesting. You said the clinician reason releases each one. This is a really cool example. I asked you for a cool use case. You gave me one. But AI kind of got popular with the promise of saving time and effort and reducing burden on clinicians. Now I've spoken to other executives and they say, oh, ambient listening has

Testa, Paul   8:01
Yeah.
Absolutely.

Jordan Cooper   8:21
actually increased provider time in the EHR, not reduced pajama time. I see you shaking your head. You can speak to that in a moment, but more specifically in this use case, it sounds like you're adding additional tasks to the already busy provider's plate. How are they responding and how is that affecting patient care?

Testa, Paul   8:39
Absolutely we are, but that is because we are in a validation phase. This is a proof of concept.

Jordan Cooper   8:42
OK.

Testa, Paul   8:45
And as I said before, innovation is not interesting if I can't scale it. So what we're doing is we're looking at the delta, the distance and edit, and we're building three kinds of judgment models that will look at those, and then we will get the human out of the loop. This is summarization. This is actually one of the more easy tasks. This is extraction and summarization.

Jordan Cooper   8:54
Mhm.

Testa, Paul   9:04
So in the validation phase, you find good clinical partners who are willing to work with you and yet put in some effort. We're going to ask something of them because they appreciate in a longer run, the patient and clinician experience will be better when we move to a judge, an AI judge-based model that extracts the clinician from it. I am, we rely too much on saying, oh, but there's a human in the loop.

Jordan Cooper   9:13
Mhm.

Testa, Paul   9:26
We find that that adds bias. It engenders inequity. It will make something less empathic. It will make something, it will often add into the reading level when we're exactly trying to lower the reading level. So I will tell you, I firmly believe always having a human in the loop is not as comforting as we think it is.

Jordan Cooper   9:45
Yeah, so you talk about clinical workflows and about how we'll be eventually removing humans from the loop in order to enable AI-driven workflows. Fantastic.

Testa, Paul   9:55
But, but eventual is like 6 weeks from now, not six years.
Oh, like the world, yeah, yeah, yeah, yeah, the world's fairs, yeah.

Jordan Cooper   10:16
But there's the other side of the coin. As AI becomes part of these clinical workflows, an intrinsic inherent part of these workflows, what does downtime mean? And how has that changed how you think about resilience?

Testa, Paul   10:30
Um...
It is part and parcel of our job to ensure enterprise resilience, right? That's an easy thing to say out loud. But there are enough proactive monitoring tools. And we also have to acknowledge where these tools sit. So can a daily summary go down and we keep the lights on?

Jordan Cooper   10:47
Mhm.

Testa, Paul   10:52
provide the number one quality care in the country. Yes, these are important systems right now, but we can all envision a time when they make or break the quality care we provide. And that's why we got to build these systems and Hardin them now. Absolutely agree with you. That said, a vast majority of these. So for example, when we have them in the

Jordan Cooper   10:59
Mm.

Testa, Paul   11:13
education environment or the research environment, the expectation of resilience and uptime is frankly lower and our doors can stay opened if we have a downtime in a place that we, in, for example, in a researcher education, we cannot have that.
In clinical care, and we measure those, we measure those times in seconds and minutes, not hours, for a year over.

Jordan Cooper   11:42
So do you have kind of paper workarounds for if you have been subjected to a cyber attack or is it more like an independent recovery environment? Or is it, I mean, I know the answer is always all of the above, but can you elaborate on what hardening your system to ensure resilience in a clinical care delivery
Setting looks like.

Testa, Paul   12:02
Yeah. So an IRE is absolutely core to our strategies. But let's also acknowledge that we've been digital for 15 years, and we know how to thoughtfully take the system down and bring it back up in a matter of, you know, it used to take 6 hours.

Jordan Cooper   12:15
Mhm.

Testa, Paul   12:22
and now takes 35 minutes. But those 35 minutes, there are people who are much more informed than I that can speak to the technical aspects. I will tell you from the bedside.
I need a clinical population. I need peers of nurses and physicians that say to me, I know why you're doing this. And it's not just to make our night on a Saturday night in June worse. That, oh, good things are coming out of this change. And that's actually an important cultural thing that they don't view.
change as a bad thing, but that we knew about it, it comes down, and when we didn't know about it, and it's an unexpected downtime, which fortunately are quite rare, we rally the same way we would rally with the same playbook at a threat to the enterprise.

Jordan Cooper   13:14
Just to be clear, and I think everyone listening knows the answer to this, but do you have the equivalent of a third grade elementary school fire drill? Meaning, do you guys do have these planned downtimes? And you're saying that these nurses and physicians now understand why they have the inconvenience of a planned downtime. That's what you're saying, correct?

Testa, Paul   13:34
Yes, thank you for taking that, yes.

Jordan Cooper   13:35
No.
OK, so...
All right, so AI is helping facilitate care. We have plans for when there is downtime, hardening our resiliency. But how has AI changed your threat model? Do you ever feel like you benefit from outside expertise? Obviously, AI is generating more frequent attacks, sometimes novel kinds of attacks, including impersonation of CEOs of vendors who you think you're talking to, but you're not.
You know, how has AI changed your threat model?

Testa, Paul   14:08
Lots of secret sauce, but we're all facing it, so it's important to be transparent with each other in the right venues. I am learning more from our CISO, our Chief Information Security Officer, in the last year than I have, probably in 15 years in this job. And he's an important partner to me. I would say we are

Jordan Cooper   14:23
Mm.

Testa, Paul   14:28
You know, we are finding...
vectors of attack that have made us close doors regularly. And it's, they're getting really creative and we have to ensure that we're using the same models and the same weapons.
to prevent these attacks.

Jordan Cooper   14:46
Mmh.
So it's kind of like in the Cold War. I have 15 nukes, I have 20 nukes, I have 25 nukes. It's an arms race in a way.

Testa, Paul   14:56
In a way it is.

Jordan Cooper   14:58
As AI usage grows, how do you decide what runs in the cloud versus on-prem and how are you keeping those costs predictable?

Testa, Paul   15:05
Ohh.
Great, great question. My pause is I don't know if I have the firm answer yet and we are learning. I think FinOps is an entire field unto itself of experts who we have employed in that situation. That said.
The ability to remain flexible and thoughtful in how, for example, if we have 40 million images, clinical images stored, and how do we keep them in deep storage, but still have some latency acceptable? Or as we move to all digital pathology,
We have millions of pathology images. The reality is we don't need instantaneous access to those. So we have to engage with the vendors to talk about what does it mean for deep storage and what do we get to put there and not always be extracting and then store locally as well. It's an absolute hybrid model. And we've

Jordan Cooper   15:48
Mhm.

Testa, Paul   16:03
Been, I think, fortunate to not commit too early to a hosted environment, but really maintain a hybridized model of on-prem and all three large hosting vendors.

Jordan Cooper   16:17
Got it. You mentioned FinOps is a field of experts, and of course, everything is a field of experts. What is hardest to hire for right now? There's a lot of AI experts, a lot of demand, probably more demand than supply of the best kind of experts that you need.
for anything from cybersecurity to developing AI models to doing your FinOps, et cetera. How do you decide what to build in-house, what expertise to build institutional knowledge in versus what to partner on?

Testa, Paul   16:50
Yeah, there's actually probably a couple of questions in there, I think.
I, we're always doing the calculus of what to build, what to buy, what to partner on. There are, I think we are, there are certain things that we've come to realize are simply commodities in this, and we have enough teams that we can build our own tools.
that there's just a level of...
inappropriate pricing. We're not a large healthcare, we're not a large financial institution. We're a nonprofit healthcare. And we have to be incredibly responsible with that margin. To do that means building things locally, but also an incredibly important strategic relationship with our EHR vendor, Epic.
and they are responsive to things we need. And they, but why are they responsive to things we need? Why do they listen? Because we will deploy at scale. I can't have five different dietary management tools. I need one. I can't have 5 instances of my EHR. I need one.

Jordan Cooper   17:57
Mhm.

Testa, Paul   18:00
If the same goes across all from our tripartite mission, it goes across all three environments of the education, research, and clinical that we should be based, we should use to the fullest that which we are already paying for. And when we are at the margin, think about whether we need to build it ourselves.
which is increasingly becoming viable, or partner with either upstarts, we find a lot of those. Those are the back hauls at HIMS and the smaller booths. We find partners early. And if they're willing to try to scale, meaning deploy everywhere for us, then it's a good partner for us.

Jordan Cooper   18:40
So I just want to ask a fun question because we're approaching the end of this podcast episode. You mentioned Ultraviolet, which I believe is the name of NYU Langone's own LLM. And then you mentioned something that is quite obvious to every listener here, which is your nonprofit health system with slim margins. Everybody can identify with that.
So something I think that's interesting is that different health systems across the country are saying, maybe we're not only in the business of providing clinical care. Given your background in business and JD and your MPH, I said, maybe, you know, you're bringing something to the table, which is maybe there are new revenue streams. Maybe we can also be a software company.
Maybe we can license out Ultraviolet to other institutions that don't own their own GPU stack in their data center from pre-Gen AI days. And maybe we can take these resources and we can be the licenser and allow others to buy from us. And now we can have 30, 40% software margins
that helps enable us to do our care delivery mission. What are your thoughts? And it's just kind of a fun, oddball question. What are your thoughts on NYU being more than just a care provider, but also a software company?

Testa, Paul   19:59
I have very strong thoughts about that. And this comes up not infrequently because I think we're good at what we do. But what we do is provide number one quality care and the safest care of the health in the country. That's Visient ranking and that's important to us. I think there are health systems that are able to or wish to live in that space

Jordan Cooper   20:13
Mm.

Testa, Paul   20:19
of being a software shop and maybe a SAS shop as well. I know my mission is to provide high quality, safe, high patient experience care. So I want to be candid that I feel at times that can be a little bit of a distraction from the core mission.
We have to generate the next generation of clinicians, of nurses and doctors. We got to do cutting edge research that will change the way we provide the care that we provide and do that in a learning health system loop where we learn from every patient. So I think we're happy to engage when we can in that sort of co-development.
But that's not where our secret sauce lies. High quality care with a great experience.

Jordan Cooper   21:07
So Paul, I'd like to turn it over to you with the final question of this interview. What advice would you give to yourself or anybody listening, but yourself maybe a year or two ago, or anybody else listening across the country who's a little bit earlier in their AI data center journey than clearly NYU Langone is?

Testa, Paul   21:32
Take the bet.
that this is not about making yourself ready for some pivot down the road, but commit early and be able to scale. And I really want to emphasize that last part. There are times we, not just me, but more senior than me folks here, we've had to make that decision of saying no to certain things.
where, because it won't scale. And I want to reiterate what we opened with. Innovation that doesn't scale, using things we have already paying for at scale is the fastest way for us to protect our margin and to be responsible with our resources. So use with what we're paying for and stop paying for things
Three times over.

Jordan Cooper   22:20
I appreciate all of your insights. We've covered a lot of ground today. For our listeners, this has been Dr. Paul Testa of the NYU Langone Health System, the CHIO. Paul, thank you so much for joining us today.

Testa, Paul   22:33
Thank you so much.

Jordan Cooper stopped transcription