#

View All Sessions

Advancing Data Quality in Public Health Through AI: The Datapult ELR Use Case

March 14, 2024

Share This Page

Speakers:

Patina Zarcone-Gagne, Director of Datapult, an APHL Company; Sarah Brumley, MHA, Client Success Manager, Director of Federal & Public Health Accounts, Clinical Architecture

This presentation highlights how Clinical Architecture worked with Datapult, an APHL subsidiary to address data quality challenges in electronic lab reporting (ELR) in public health. They discuss the expanded ELR that was created and discuss the vision for the future – seamless data exchange during public health emergencies and the use of quality data to stop the spread of infectious diseases.
h

View Transcript

Transcript

View Transcript

Dr. Victor Lee (00:04):
Good morning everyone and Happy Pie Day, March 14th, 3.14. For those of you who celebrate, I certainly do. Good morning, my name is Victor Lee. I am Vice President of Clinical Informatics at Clinical Architecture and I am pleased to be moderating our first session today about our efforts with the Association of Public Health Laboratories on an electronic laboratory reporting use case. I am joined by two excellent speakers here. Sarah Brumley has a Master’s in Healthcare Administration and is our Client Success Manager at Clinical Architecture. She is also the Director of Federal and Public Health accounts and she’s responsible for our relationship with the Association of Public Health Labs. We also are joined by Patina Zarcone. She has a masters in Public Health. She has a wealth of experience in informatics and she is the director of Datapult, which is a company under the Association of Public Health Laboratories and they’ll be sharing some fantastic insights about what we’ve done together in partnership to advance an electronic lab reporting use case. These are their photos and their background, and so we’ll explore how we partner together to expand an electronic lab reporting use case and Patina, I’ll hand it over to you for starters.

Patina Zarcone-Gagne (01:50):
Wonderful, thank you. Thank you so much, Victor. I will advance to the next slide. So I’m going to briefly introduce Datapult to you guys. As Victor mentioned, Datapult is a new company that is a wholly owned subsidiary of our mother company, APHL. We are a nonprofit and we are located in Bethesda, Maryland, although our staff are located all over the country. APHL is a membership association for government laboratories that do testing of public health significance. We support local, state and federal organizations with the range of issues in the area of informatics and public health.

(02:40):
So that’s who we are. Now, what do we do? We have three different main offerings: messaging services, data platform and support services. Our messaging services is really where our expertise is. As you can see on the map, our cloud platform is connected to every state in the country as well as the territories and some of the larger locals like New York City, Houston, LA, and we have a variety of different, does this have a laser pointer? No, I don’t think so. Okay, I’ll step over here. We have a variety of different use cases that go on within the environment. The one that we’re going to talk about specifically today is expanded ELR, but you can see we do, we send data for infectious diseases and VPD, vaccine preventable diseases, we do electronic test order and result reporting in our environment. We have advanced molecular detection and genetics data going through. We are also a platform as a service, so we do host applications sometimes on behalf of other public health agencies and some are applications for services that we provide. And we also have an informatics consulting arm. So we provide a lot of consulting services, custom consulting services to states and locals, public health departments.

(04:19):
So movements of public health data, as I said earlier, is our expertise. We have over 250 data partners utilizing our cloud every day and we have hundreds of terabytes of data that are sent, lab data, that are sent across our platform every month. So with this you can see how important quality lab data is, and we’re going to go into that a little bit later. I’m going to pass it over to Sarah to introduce Clinical Architecture.

Sarah Brumley, MHA (04:50):
Alright, so Clinical Architecture, we were founded in 2007. We offer a comprehensive terminology management solution focusing on enterprise interoperability from a semantic and syntactic perspective. We were founded by a group of industry leaders that have a wealth of knowledge in the space and our customers really represent on the right hand side there. We are really across the healthcare spectrum. Everybody has terminology needs, data needs, being able to utilize that data. It’s not a matter of having more data, it’s having better data. And so health systems, government space, analytics vendors, payers, we help all of them to get more benefit from their data. So some of the things that we do, terminology management, like I said, semantic/syntactic interoperability, clinical inferencing, which is something that we will speak more to here later in the slides. Value set authoring and normalization. So very quickly who we are.

Patina Zarcone-Gagne (06:03):
Thank you, Sarah. So what is the current state of data quality as we see it across public health? So factors influencing data quality, we’ve bucketed them into four main areas. We’ve got workforce policies, processes, and technology, and there are challenges in all of these four buckets. So let’s talk about workforce first and what I’m describing here is within the public health. So we’re looking at the public health system specifically state health departments. So training limitations are just really a big issue. There’s not enough training available and the training that is available is very expensive. So it’s hard for them to get trained in some of these very particular expertise areas. Available salaries and turnover. We have folks coming into public health, they get trained, they do get trained, they learn and then they go on to private industry where they can get paid more and it happens, the turnover happens pretty quickly.

(07:20):
And also the expertise required. They’re just lacking this expertise. They need to have more knowledge and data standards to be able to do a lot of this work. And it’s just the capacity isn’t there. So let’s talk about policies next. Reporting mandates vary by state, like what actually is reportable, what isn’t reportable, what the state wants, what the state doesn’t want. So there’s a lot of variability there. Communication of reporting requirements is inconsistent. So I know Victor and my colleagues here at Clinical Architecture can say something about that, but there are siloed systems and there’s data everywhere. And so locating all of that data to be able to get it into one system or get it moving with data standards and increasing the quality is very difficult. And a lot of times our colleagues here have had to deal with free text, which is a big challenge. And there are also outdated state technology policies. Some of our states can’t even use cloud, and if they can use cloud, sometimes it’s a very specific cloud. They have to use Google but we are on Amazon and so there are a lot of different technical policies. I know I don’t want to paint such a sad picture here early in the morning, sorry guys, but these really are very big challenges. In order for us to move forward, we need to get past some of these and find creative ways to do that.

(09:05):
Other factors influencing data quality processes. There are too many interfaces to maintain, going point to point in connections. I mean, I don’t need to probably tell you guys how difficult that is to do and maintain all of those different point to point connections. Manual review and feedback. Not only is this manual review, which takes a lot of time away from people that are doing five jobs, but also you need to know how to do it, and there’s also an element, of course, of human error in doing that. All relevant patient data is not collected. This is a big issue. Some of the laboratories that need to report to public health departments, they’re just not getting enough of the demographic data to be able to have that particular data set be actionable by public health. And in addition, maximum number of concurrent onboardings. Our states are very spread thin and so onboarding different laboratories, for example, to be able to send this data for the state to get it,

(10:20):
it takes, I mean some of our states have a backlog of nine months. So it is really, really a challenge. And then I’ll talk about technology a little bit. So again, I mentioned earlier, manual data entry and human error is just going to happen, right? Systems may not produce HL7 natively. So if that’s the case, maybe it comes out in a CSV, but our state health departments require HL7 ELR specification right now. So how do we help those where HL7 does not come out natively. Siloed systems, as I mentioned earlier too, a lot of these systems just aren’t speaking to one another. And so that really does decrease data quality and modernizing infrastructure is costly. Getting funding for states and for locals right now there’s money out there, but it’s feast or famine and sometimes you get a lot of money. Our states get a lot of money, they’re not sure what to do with it and then we help them do with it, but then the question is, but how am I going to sustain this? Because this money is going to go away. So we are going to, it all goes up from here. Sarah’s going to bring us and show us how we’ve done work that can solve some of these data quality issues and help with capacity and capability. So with that, I’ll turn it over to Sarah.

Sarah Brumley, MHA (12:00):
Alright, so as Patina mentioned, it’s mandated to report these notifiable conditions in the ELR specifications. So what we wanted to do is set the stage for all of the factors that Patina went through. What does it look like today? The data in a lab, given all of those factors. So this was a true sample message that was pulled. Obviously no PHI, everything is masked, but this is what the quality of data looks like in a sample HL7 message that would in theory be used as a reportable to a jurisdiction that is going to be receiving it. For that one HL7 message, there are 54 corrections that need to be made in order for it to be ELR compliant.

(12:55):
And this is standard, this is not an outlier, this is not something that was fabricated to be the messiest message that we could find. This is the true state of data quality in labs today, and it is just what is essentially the minimum that can be done to do what is required, but then you have these extra layers on top of it. HL7 is just a guideline, but then you have ELR compliance on top of that that is required for these reportable conditions. So you can see that some invalid characters in the zip code. Some data is just not found at all and these are all things that have to be done. So this is the then back and forth that occurs between the labs and the receiving jurisdictions to be able to get those messages reported. On top of that, you have varying reportability criteria, and this is what Patina had brought up in the beginning is it varies from state to state or in some cases you have several jurisdictions within one state and they may have different reportability criteria.

(14:12):
In that case, say you are a nationwide lab, you have several sites, and then that means you have several jurisdictions that you have to report to. Here’s just an example of three for one reportable condition. There are up to 90 reportable conditions. And if you have, for example, we have one lab that has 17 states or they have physical locations in 17 states and they have 25 states they have to adhere these requirements to. So just think about the scale of that, all of their conditions, all of these reportability criteria by state and somebody has to manually track all of that and these change based on when the state is requiring them to be changed. So somebody has to keep track of all of that and make sure that you’re adhering to all the ELR criteria. You’ve got all of your connections set up and everything has been vetted to say, yep, this looks good.

(15:18):
So with all of this in mind, Datapult partnered with Clinical Architecture to create expanded ELR. Really what this is intended to do is to improve the data quality in public health and alleviate a lot of the burden on the labs or the sites that have reportable conditions that are sending those to the jurisdictions. What we did is created a custom solution using Symedical, Pivot and the Clinical Awareness Suite to semantically and syntactically transform those messages. So currently we are taking in HL7 messages and outputting ELR compliant messages and it’s available on 90+ conditions. So when labs come forward with here’s our test menu, here’s what’s reportable, we’ve got those set up, we create their own essentially set of those reportable conditions. This also provides a one-to-many connection. So those labs will just connect with Datapult and then everything just gets delivered through.

(16:23):
We’ll walk through the data flow later in just a few slides here, but instead of maintaining, what, 25 connections to those separate jurisdictions, setting those up and then maintaining them, it’s just a one to many. So alleviating a lot of that burden there. The Clinical Awareness Suite is what determines what’s reportable as well as identifying the jurisdiction that those messages are going to be sent to. A notifiable condition just so that everybody’s on the same page, that is what labs are required to report to the jurisdiction. So as I’ve mentioned before, these are often states and those states say what is reportable as well as the specific criteria within that reportability. And as I mentioned, it can change, new conditions can be added.

(17:16):
So here’s our data flow. So it all starts with those labs sending in the messages to AIMS and that’s in a version 2.x. It comes into Datapult, they push it to us. Pivot, which is doing the format transform, takes those from version 2.x to ELR compliant messages. We then run through all of the mapping and the term validations. We confirm that it is conformant. If for some reason it is not, then we will quarantine that and it’s not going to go out to the jurisdiction without being truly ELR compliant. So we don’t need to worry about messages going out that do not meet that standard. And then those are addressed via human intervention. If it is ELR compliant, we’ll look to see does this message have a code in it that is present in our inferences or in the values that we would expect to find within those conditions?

(18:17):
If yes, then we’ll determine which jurisdiction. So we had said everybody has their own flavors and factors within those conditions of how they want to receive those messages. So we have individual inference sets essentially built out for each jurisdiction for exactly how they want to receive that information. And it is quite a lengthy process to go through and work with each jurisdiction to figure out exactly what they want. So we determine which set it should run against, run those inferences on it to say, here’s my message, is this something that is reportable? Is it not? Should it get sent on? Is it meeting the criteria? If all of the above are true, we send it back, we construct the message, send it back to AIMS, and then it is passed out to the jurisdiction. So it’s really taking away a lot of the effort. Because there is a lot of upfront work to getting this set up and we have gone through that process, but once you get that initial lift, it is pretty streamlined after that point.

Patina Zarcone-Gagne (19:25):
Yeah, it really is. And to your point, Sarah, that initial lift was astronomical. I mean, this was never done before. This doesn’t exist. This exists for things like electronic case reporting, for example, but this is electronic lab reporting and this logic engine is very unique in the sense that yes, we gathered every reportability rule from every single state. We started with their websites, a lot of them have that information. So we pulled that down and then we would have to go back to the state and say, is this still right? Is this up to date? Are there things that you’re changing? Every single state in the country. And that data did not exist in one central location. So now it does.

Dr. Victor Lee (20:24):
I thought I’d also take a second, but maybe go back to that slide just to connect the dots with something that Sarah mentioned about the 54 corrections in that HL7 message. And you use the example of an invalid zip code. What’s the consequence of not having a valid zip code within the context of this data flow?

Sarah Brumley, MHA (20:44):
Yes. If you don’t have a valid zip code, then we can’t identify what the jurisdictions are. So each jurisdiction is essentially identified by their full list of zips. And so when we have that incoming message, we use that zip code to identify which inference set to run on that message.

Dr. Victor Lee (21:06):
And that’s just one of the 54 corrections. So there’s a tremendous amount of data quality improvement opportunity.

Sarah Brumley, MHA (21:16):
Alright, so overall what we’re trying to do here is, I think Patina had mentioned just how burdensome and tedious this can be for the actual workers that are in the labs in public health. So the big points to take away would be we are really alleviating a lot of that burden. It is a connection to one site rather than multiple jurisdictions. As a reference point, I think that you had said that one of the backlogs was nine months to establish that connection. But even at that, once you start the process, each one can take up to 12 weeks. You have 25 of those plus all of the additional effort that we’ve already talked about. I mean it just keeps piling up.

Patina Zarcone-Gagne (22:05):
Plus the day job, yes, at these places.

Sarah Brumley, MHA (22:08):
Yes. So in addition to that interface development setup and maintenance, that is essentially off the plate when it comes to the electronic lab reporting aspect.

(22:23):
Obviously there’s going to be other inferences and maintenance needed on other projects, but when it comes to this, that is essentially not a factor anymore. And then just reducing that churn for the validation for the ELR compliance. So once we go through the process, we get those sample messages, we identify the 54, or however many it is, areas where those messages need to be improved, then that lab is set. They’re not going to have to keep going through that process with all of the other jurisdictions that they’re going to be connecting to because they’ve already done it and we’ve confirmed it. We’re running it through all the same software again. So in addition, all of the data is going to be mapped to the standards Patina had referenced that the expertise in these sites, it’s just, I think it was said best when I had heard it earlier today.

(23:13):
Public health is just dangerously underfunded in a lot of ways and getting that expertise in is such a necessity in any way that we can. We can do it through this and better data equals better outcomes. So having that data mapped and having the quality in those messages and then getting it to the public health is huge. And then alleviating the manual effort for aggregating and updating jurisdiction requirements, more comprehensive and accurate data to those public health sites enables improved population health management. So really just improving data quality, improving the population health quality, the ability to make actionable decisions. So that’s what we’re doing there. I don’t know, do you want to talk any about where you envision the future state to be going?

Patina Zarcone-Gagne (24:07):
Yeah, and I guess just one, the last slide gave me some ideas here and thinking about the pandemic and the beginning of the pandemic, people were running around, we need this data, all the way up to the federal government, we need this data. And we were fractured, really, really siloed, non interoperable. And I believe that what we’re doing here is going to be one of those things that if this happens again, we’re going to have a model by which we can get this data fast, accurate, and in a standards-based format to those that need it. And so that’s part of the future vision. A little segue, CSV to HL7 to further alleviate burden on labs for reporting requirements. We can do that. We can take data in CSV, run it through the system, as Sarah mentioned, and turn that into the compliant message that our public health departments need. Public health system that has information it needs to conduct targeted and efficient interventions within their communities.

(25:22):
I mean, that’s what we all want. And if we can get the data and we can look at the data and we can get into the hands of our state departments, that’s something that’s very, very important to do. Quality data that shows disparities in healthcare and can use it to intervene and shore up gaps with more data and with better data we can get at this data quality and disparities issues and hopefully find some solutions. Systems in place for seamless exchange of data in times of public health emergencies. Hence the COVID story, pandemic. Quality data to stop the spread of infectious diseases. That’s also something we all want. We can have case reports, we can have presumptive, I think you have COVID, but the lab data is what diagnoses it. And so we need that data in order to stop the spread of new emerging infectious diseases. And all this ultimately improves the lives of populations of people in the United States and globally. So I think we’re at the end.

Dr. Victor Lee (26:40):
Great. Thank you so much.

Patina Zarcone-Gagne (26:41):
Thank you.

Dr. Victor Lee (26:42):
Sarah, I’m going to open it up for questions in just a second, but first I have a comment and then a question. I’ve had the opportunity to work with you all on this solution, and I just want to say that from an informatics perspective, this has been one of the most rewarding projects that I’ve had an opportunity to work on because of all the informatics principles that are tied together in order to make this entire solution work. So I’m very appreciative of that opportunity. The question I have is for you, Patina, if there are any laboratories out there, this audience or watching the recording and they want to connect to this ecosystem, what should they do?

Patina Zarcone-Gagne (27:24):
Well, what they would do is contact either Sarah or ourselves we can be in contact with. We work together. And so I should have had my information up here, not a QR code for us, but I can, forgot that minor detail, sorry. But Datapult, we have our website, it’s datapultaphl.org. Please come visit. There’s contact information on that for us. We have a phone number, we have a QR code there, we’ve got all the good stuff. So it’s datapultaphl.org. Thank you for asking that.

Dr. Victor Lee (28:07):
Great.

Patina Zarcone-Gagne (28:07):
And Victor, I want to say you’ve been wonderful in working with and in this project, the technical complexity that you had to overcome in some of this was really quite amazing. So thank you.

Dr. Victor Lee (28:20):
That was a pleasure. Okay, thank you. Any questions from the audience? Okay.

Audience Speaker 1 (28:27):
Thanks, Joe Bormel, I’m a board certified internist. I’ve done some work for the CDC and I read Scott Gottlieb’s Uncontrolled Spread. I’m interested in how close are we to getting to sequencing every patient so that we can see the extent of types of things like COVID are and give me a sense, because you’ve got national coverage and also you serve state labs.

Patina Zarcone-Gagne (28:58):
Right. So I will attempt to answer this. I think I know a couple inches deep, so with the CDC data modernization funding to our labs, of course there was a request for the states to beef up their infrastructure around next generation sequencing and advanced molecular detection. So that is actually, I mean, and I’ll speak from the standpoint of our public health laboratories because that’s, and for those of you who don’t know, every state has a public health laboratory at the state level. Some of them have laboratories even at the local level, cities and counties. So when I reference this, I’m really talking about the public health labs in each state. They are utilizing that money to ramp up that infrastructure.

(29:55):
The problem is they can do the sequencing, but what do they do with that data after it comes off, right? One of the things I didn’t mention here is Datapult on our cloud platform, we actually have partnered with Seqera Labs. They have Nextflow, I’m not sure if any of you, it’s a pipeline management software. And so we help them. We actually have that in our environment. But what we did was we built on top of that a graphic user interface where you don’t have to be a bioinformatician. Our states, many, many don’t have a bioinformatician, and so they don’t have the working knowledge to use these complex systems like what you’re just talking about. So we made an easy button to put on top, basically you don’t have to be a bioinformatician, it’s preloaded with pipelines, and that’s something we’re just trying to make it easy. And so I mean all 50 states are working in this in some capacity. Thank you for the question.

Dr. Victor Lee (31:05):
Any other questions from the audience? Yes, right over here.

Audience Speaker 2 (31:11):
Good morning, my name is Azada. Question. You kept saying if we get the data we get.

(31:17):
You kept saying if we get the data, if we get the data, what do you mean by if we get the data? And the second question, how you deal with the military. Okay. By the way, I’m from Public Health Department of Defense and we are the center of excellence for HL7. So I just wanted to make sure that you’re aware of it before you answer my question.

Patina Zarcone-Gagne (31:40):
Sure, sure. And Victor, I’m so sorry. I apologize. I have a little hearing loss on this ear. So can you rephrase the question?

Dr. Victor Lee (31:48):
I couldn’t entirely here it either.

Patina Zarcone-Gagne (31:49):
Can I get closer?

Audience Speaker 2 (31:51):
Sure.

Patina Zarcone-Gagne (31:51):
Let me come over.

Audience Speaker 2 (31:52):
My first question, you mentioned that if you get the data, if you receive the data, what do you mean by that? So look like there’s a way of maybe we will not receive the data. And the second question, how you deal with the military, like all the ships and all the carriers or other military treatment facilities and if you deal with it or not, and you mentioned something about the zip code, why we have problem with zip code if we get the lab data and all the HL7 are from hospital or facility because you’re going to relate that to the facility itself and you get the zip code except on the individual cases, which is like they reported for certain type of disease. And that’s where you get the case finding from. And I’m from public health through DHA.

Patina Zarcone-Gagne (32:54):
Awesome.

Audience Speaker 2 (32:54):
And we are, again, IRP, the center of excellence for HL7, but HL7 for the military is different from the civilians. And we report to CDC if anything happened that they don’t know about, which is using veers. If any case come out, we have to report it. So if I know have too many questions, but if you can answer any of them, that’s fine.

Patina Zarcone-Gagne (33:19):
Okay, thank you. Thank you. No, I love it. We love questions. That’s why we’re here. I’m going to attempt to, if I didn’t understand something, please forgive me and correct me. So I’m going to start from the hospital question and the demographic information. I mean, you guys probably have more experience talking about that than I because I would assume, I’m assuming we actually Datapult itself just to date has worked with large diagnostic clinical laboratories. Like, for example, a Quest, or for example, a LabCorp. Just as examples, we work with clinical labs. We haven’t actually had the opportunity yet to connect to a large hospital system. So we’re working on that. We’re only three years young, so we’re still infants in this area, but we are growing and hospital systems is definitely one that we want to target. I mean, I am making an assumption here, so don’t quote me, but I would say hospitals still don’t get demographic information sometimes. I mean, I would assume there are people, maybe transient people that come in, they don’t have demographic information. These are just some ideas I’m putting out there that the record is never a hundred percent complete or at least I haven’t seen that from a lot of these different organizations that we work with.

Audience Speaker 2 (35:10):
There’s other databases you can link to get the demographic data to be replaced. And also there is situation where the medical facility, there is admitting facility which have occurred and the treatment facility, what does that mean? Facility have admitted the patient, but they have to transfer that patient to another facility to get treatment. So now you have two records and which record you going to use. So we use some other databases to collect the demographic data if we don’t have it.

Patina Zarcone-Gagne (35:42):
Got it, got it. And I’ll be honest, we haven’t really gotten that far into the hospital systems yet. I mean the closest example I can say to this is that someone gets sick in Florida, they have a lab test right, here, but they’re not from Florida. They live in Massachusetts. And so, correct me if I’m wrong, our system is able to send that data both to where the lab test was taken in Florida and also to their home state in Massachusetts. So that’s a little bit of that, I know you’re talking much more. Yep, yep, yep. And as far as working with the military, we have worked with the military. We did during the pandemic, the Department of Health, DHA Department of Health Administration, we sent their COVID data to public health. So we’ve had the opportunity to work with them. We haven’t had the opportunity, well, we also have had the opportunity to do the same for the Veterans Administration. And we work actually another federal agency, we work with this NIH and we do over the counter test result reporting for them. So some of that’s still going on, but that’s the extent I’d love to have the Defense Health Agency look at this expansion and see if they’re interested in coming and partnering with us again. That would be great.

Audience Speaker 2 (37:26):
For example, the Roosevelt and before it went to anybody else, it came to us as an application to collect the data. So it was way ahead of everybody because we have to deal in real life situation moment where you see it, we create the case finding and send the result to people there and facilities that need it. That’s why we kept them somewhere else and track everything with their movement and we designed the system to do that. But I appreciate your answer. That’s good.

Patina Zarcone-Gagne (38:02):
Thank you for your questions. They’re very insightful.

Dr. Victor Lee (38:07):
Okay. I think we’re about at the end of our time slot. Thank you so much, Patina and Sarah for sharing these wonderful insights and as well as the audience for the excellent questions. Thank you very much and have a great day.