View All Sessions
The Use of Data Models and Terminology in Analytics
March 13, 2024
Speakers:
Phil Lindemann, Vice President for Data and Analytics Informatics, Epic; Suzy Roy, Customer Relations Manager, Americas & Collaborations Specialist, SNOMED International; Ben Hamlin, DrPH, Senior Research Informaticist, NCQA
View Transcript
Transcript
View Transcript
Carol Macumber, MS, PMP, FAMIA (00:03):
Good morning. Thank you for welcoming. Thank you for joining us today for this presentation on the use of data models and terminology for analytics. I’m joined here with some esteemed colleagues. This time I’m going to remember to change the slide as we start. This is me, I’m Carol Macumber. I’m the EVP of Client Services at Clinical Architecture. Today, the topic is around terminology, models and terminology and analytics. We’re going to talk about the importance of those common data models and standard terminology for purposes of analytics along with its real-world application in an electronic health record platform. Joining me today, I am pleased to welcome my esteemed colleagues. Dr. Benjamin Hamlin from the National Committee for Quality Assurance. He is a distinguished senior research informaticist specializing in clinical quality person-centered clinical decision support and predictive analytics renowned for leadership in digital health and quality measurement. Dr. Hamlin seamlessly integrates academic rigor with innovative measurement solutions as a national leader in transformative quality strategies, he strategically leverages health information technology, advocating for human-centered design and cognitive theorems in healthcare quality evolution. Welcome Dr. Hamlin.
(01:25):
Suzy Roy from SNOMED International has been involved in the health data standard space for the past decade from creating and maintaining standards to providing strategic insight and collaborative opportunities for governments and organizations. She serves as the Customer Relations Manager for the Americas, the Collaboration Specialist and the Research Engagement Lead at SNOMED International. As part of her work, in this role, Suzy thrives on engaging with stakeholders and promoting the use of SNOMED CT for National Health IT strategy and research. She’s always seeking to facilitate the links for harmonization and interoperability between standards. Suzy holds two master’s degrees, one in experimental psychology and another in library information science along with PhD work in behavioral neuroscience. Last but certainly not least, Phil Lindemann is the VP of Data and Analytics for Epic, our giant neighbor. He works with all of Epic’s software applications to help them use data to better inform product design and roadmap. Phil is experienced in creating executive analytics, embedding machine learning and clinical care, and designing systems for multi-institutional clinical research. His current focus is Cosmos, a software application that allows collaboration between hundreds of health systems to advance medicine through rapid research and point-of-care tools. Phil graduated from the University of Wisconsin and lives in Madison.
(02:52):
So we are going to start with Suzy talking about terminology as a precursor for interoperability and analytics.
Suzy Roy (03:03):
So hi everyone. Yes, I’m from SNOMED International and I’m here to talk about SNOMED, but also I am going to continue to point out where and how utilizing international standards is actually going to assist you in conducting your clinical data analytics and further research and how that’s going to feed into information models and then ultimately how you’re going to be able to do your analytics. So just a really brief primer on SNOMED CT itself. So SNOMED CT is the world’s largest clinical terminology. It is utilized in electronic health records and data health systems around the world. And again, it’s not just for electronic health records but also for research records and data systems as well. With SNOMED CT, you have concepts which is essentially a clinical idea, and those are provided with a unique identifier and those are linked to other concepts by a relationship.
(04:10):
And so really this is how you start to build an information knowledge model utilizing the terminology itself when you have a concept that is in another type of concept. So in here you have an excision of an appendix procedure is in a type of procedure, but with SNOMED CT, we don’t just have “is a” relationships, we also have other types of relationships like it’s in a procedure site and that’s going to be in a particular body structure or you’re going to have a method. And so you can really start to see how and where this is building your knowledge base. But SNOMED CT is not just a typical linear classification, it’s actually a poly hierarchy. And so here you can just start to pick and see where in the hierarchy your concepts are going to be. But also this is where some of the power is going to start to come out. So you can do an extraction of data, of excisions, of procedures on abdomen, and that’s how you start to drill down into some of your data. And this I’ll explain in just a moment. Go ahead.
(05:23):
You can totally excise things from the abdomen. Come on. The secret sauce is really around the expression constraint language, ECL. And this is where because of that structure that I just talked about, the concepts and your relationships in SNOMED CT and how it’s formed in that poly hierarchy, you’re actually able to start to build clinical queries from your terminology. So here, just as an example, find patients with any type of viral disease. And so there is a particular way in which you can build this query. So you can start to pull out a cohort of any of your patients who have a viral disease, but in order to really be able to do this, next slide, please, Carol, you have to have some meaning or you have to have a model so that you can derive some of that context. And really it’s because of needing something that’s explicitly defined.
(06:25):
So when you have your SNOMED CT in your data, you can’t just have something that says allergy because what does that allergy mean? Is it an allergy that’s associated with family history of allergy? Is it a current diagnosis of allergy or is this a past history of the patient’s self with allergy? And so that’s where you actually start to need to utilize various information models. And with SNOMED, because of that expression language that I just showed you based on the concepts, with all of those relationships, you’re able to build specific queries so that you can start to build explicit or implicit value sets that are based on a particular information model. So just a really quick example, the International Patient Summary, which is being utilized globally right now, some of the data elements in there such as the allergy and intolerance FHIR value set based on HL7 FHIR, it is built explicitly on SNOMED CT ECL around pulling out all the SNOMED CT concepts around allergy.
(07:40):
And so you can start to see that then anyone who’s utilizing the HL7 allergy and intolerance IPS value set then has all of the concepts specific that they need from SNOMED CT for allergy and intolerance. So it doesn’t matter if someone else is also using a different sort of model, as long as they’re both utilizing allergy and intolerance, they’re going to be able to normalize those data and standardize those data because of the various information models that are for that specific, I guess clinical concept type. Okay. So essentially what I wanted to just explain is that you really need the standardized terminology under everything built into your information model, which will then allow you to do your post-secondary data analysis, whether that be for patient care and treatment or population health or research. And so hopefully my segue will now lead over to my colleague here who’ll be able to give you a little bit more information around information models and other sorts of data models.
Ben Hamlin, DrPH (08:58):
Thanks, Suzy. Okay, so I’m going to be talking today about a very specific aspect of data quality assessment, which is called Fitness for Use. And this is how we can in this area of data Fitness for Use assessment or evaluation, take a lesson from the observational research community who uses standard terminology and information models to conduct evidence-based research. And the reason I’m saying this is very relevant to both domains is because understanding the underlying essences of patterns within a dataset to understand how it’s meaningful and how it can be used and anything that exhibits its quality is essentially what observational researchers do in the pursuit of generating evidence. It’s exactly the same thing that we do when we’re doing data fitness assessment, looking for those underlying essences that have meaning within the patterns of the data that say we can use this for a very specific context of use.
(09:52):
While data quality assessment, every actor in the entire chain is responsible for the integrity of this information as it flows along into its secondary use, which is where in quality measurement in research, we consider that all secondary use of data, the end user in the quality domain is really the one who’s decreeing these parameters for expectations for data quality. It’s not that they say, I know high quality data when I see it. They have very specific context of what data quality means to them. And those user expectations are really critical in the fitness for use assessment because that’s guiding every single piece of that evaluation. So this idea of end user trust is also a central factor in data fitness assessment, right? The data fitness evaluators who are seeing these aggregate report cards, dashboards, and other specific information about a data sets existence, if you will, has to be sort of readily verifiable.
(10:44):
It has to be credible to that person and to the end user and the context for which it’s being used. And so one of the things that is really important is how do you make that transparent and much less subjective? And so the way to do that is essentially to create these frameworks and structures for allowing us to understand best practices for very specific context, which is hard because each context has its own unique requirements. However, it’s still possible to standardize those end user expectations under certain domains. And again, the observational research community has done this in their pursuit of clinical evidence for clinical guidelines for a number of rare diseases and other very specific context. So it’s all been done before, and I’m just saying here in the data quality assessment world that we should maybe take a lesson from their book and understand how to create these mappings, how to utilize terminology appropriately, et cetera, et cetera.
(11:32):
Sorry about the little lecture on data quality assessment, but this is something I think is very important. The majority of data quality assessment models are really focused on the structural aspects of data quality, the categorization at the data element level or data point level of whether this particular data element meets or does not meet expectations. It’s really focused on the structural aspects of those. It’s defined by specific standards, whether it’s FHIR or some other data model. The difference in data fitness for use is that these data models are largely ambiguous because of the large numbers of contexts for use for these and the specific requirements for each of those contexts of use. And so data fitness evaluators are really forced to or required to utilize their own individual expertise in their evaluation process. And that expertise is built up from years of looking at similar data sets under similar circumstances for similar context of use.
(12:22):
And the people who do this are very good at their jobs. It’s very subjective in many ways, and it makes it very difficult for us to standardize those expectations to then get to a place where we can scale that model up to large big data environments because the volumes of that data is just too much for any individual human to take a look at and evaluate. And so the standardization improves the consistency of decision-making by allowing us to do that. But again, because the current environment for data fitness assessment really relies on an individual’s unique expertise, we are really struggling to again move this to the next phase, which is creating computable artifacts, knowledge artifacts that allow us to help facilitate the consistency of this decision-making. So the OHDSI community is a community that has built not only A CDM, but they’ve also built a number of software tools to support that CDM as well as an entire network of researchers working with that CDM in the production of reliable evidence to inform a number of different studies.
(13:21):
It’s a very valuable tool in terms of the community support, but also the tooling that is very user-friendly. And you can go to the next slide, please. I think this my, here’s a listing of some of the software tools. So again, it’s not just the matter of having a common data model or CDM, it’s a matter of supporting that common data model with best practices built as software tools. So the best practices in the OHDSI community are actually built as software applications to apply to the common data model, which again, improves the consistency and reliability and reproducibility of all that they’re generating in terms of evidence, which is really critical, especially when you’re trying to apply that to a fitness for use domain because again, that reproducibility reliability is really kind of the fundamental underlying factor in what we’re trying to get to.
(14:03):
So the other nice thing about the community is that they’ve wrapped this all up into a single application. So the application’s called Atlas, and here’s a demonstration screen from Atlas. You can see down this left navigation bar. These are all the different software applications that are wrapped up within Atlas. So you can do all your evidence-based research here, whether it’s characterization, whether it’s incident reporting, whether it’s prediction models, all against the CDM. It’s all, again, in this one application, this is a screenshot of one queer eyes writing for a quality measure again to produce these reproducible artifacts. And then we can run these against real world data sets that have been ETL up to OMOP instance again, because that consistency will allow us to produce a lot of evidence. And you saw on the previous slide, the number of studies that are published using these tools is extraordinary. So the published literature now is generating dozens if not hundreds of studies using these methods. And so again, I think it’s just, there’s almost no argument you could make that says this wouldn’t apply to data quality assessment domain. And that’s where I stop, but I think I hand over to Phil.
Phil Lindemann (15:11):
So I’m going to take both of these pieces and put ’em together and talk about how we’ve implemented this in real life with our community health systems that use Epic. So I’m going to, I’ll just take it. All right,
(15:30):
So what we have worked with right now, it’s about 60% of Epic health systems are live with this community collaboration called Cosmos. And you probably see some signs about it in our booth, but what they’ve done is they’ve been able to bring all of their data together, de-identified and deduplicated, with data quality assessments so that they can do the types of observational research that the OHDSI community is talking about. So we go to their conferences, we learn those methodologies and then essentially implement them for our community on top of it. And it’s all underpinned by SNOMED and other ontologies like LOINC, all the things you would expect, but that doesn’t cover everything. So that’s part of what we’re going to talk about here is if you have a data model and you have a set of terminologies, what does that look like in practice?
(16:20):
So we’ve definitely got a lot of data, but when we’re talking about data quality, the most high quality data is the second it leaves the physician’s brain and goes into the computer and then from there on it’s just going to degrade and degrade. So part of the problem that we’re trying to solve is catch it upfront if we can. So we have the benefit of also being sort of the data creation system, but also the research system so we can see this full arc and all the problems that come along with it. So what we do is we use the exact same data validation checks of completeness, conformance and plausibility, but we start with completeness. So what we do, this is just the main set of data elements that are in the shared research platform and the percentage of the community that has populated that value.
(17:13):
And what we do is when a new healthcare system onboards, we look and see if they haven’t met certain thresholds and we say, why aren’t you collecting race? Why isn’t that there? And it may be like, well, we live in Lebanon and we don’t have that terminology set of the US census, so we have to work through all those things. Or there’s certain things that only apply in a pediatric setting. So this is sort of the first blush of are there data quality problems with this site? And if you have the security that I have, I can click any one of these and see the entire community and how they’re distributed along there. And this is the first step that we call green lighting. And basically if a health system’s overall data quality is not up to snuff, we will not move them into the reporting layer.
(17:56):
They can still access Cosmos, they’re still participating, but we don’t want to add it to the official layer that research is going to be done on. And there are, I think at this point, two sites that are in that where we’re working through. So it is actually an effective thing to keep some gatekeeping in there. So drilling down another layer, what we do next is we look at specific domains of data. So this is cancer staging data, and we allow an individual site to log into Cosmos and see how they can compare versus the rest of the Cosmos community on completeness, conformance and plausibility. Plausibility are the set of rules that will never be done. This is where your 160 pound babies are going to reside, right? You should never have that in your data set, but that’s like it could be filled in with a weight and that weight could be numeric.
(18:43):
So it is a complete data element. It is conformant 160 pounds is a valid weight, but it should never exist on a two month old. This plausibility list is a never ending almost creative set of data quality rules. And that’s where the difficulty is. We can make these nice clean boxes and say, okay, this is data quality, but there’s a never ending amount of metrics that we should all think about. And I really think we’re not trying to say this is proprietary. We’re happy to share this with anyone who wants to check these types of things, but it’s a never growing list. And we have this for all different domains.
(19:18):
This is an example of some of the things that we’ll catch. This was a real, real life example. I don’t know if any of you speak Afar or know anyone who speaks Afar. It’s East African language. And we had one organization onboard into Cosmos and all of a sudden 99.5% of the data elements of patients who spoke Afar were from this one organization. Now, it was a stupid error. They had mapped the language of English to Afar and then loaded their thing in here, but we auto caught this and flipped it. So what’s important about this is when we use terminologies like SNOMED and LOINC, those are the same terminologies that are used to move a patient’s record for patient care. So we’re making clinical decisions on that. When they catch an error like that, they’re actually updating what their mappings were for clinical exchange.
(20:07):
So by having everyone go onto this platform and validate their data, it actually improves the overall quality of clinical data exchange as well, not just research. So that’s really important to us to think about that. So that’s a plausibility rule that failed there. Was it plausible that 96% of the database was from one organization speaking Afar? Likely not. This is getting into the weeds, so if everyone wants to zone out for a little bit, but I think it’s actually the coolest part. So what we do when we onboard a site is there’s things that there isn’t a LOINC code, there isn’t a ICD-10 code and there’s a little bit of manual mapping where they have to do some analysis. So let’s just look at the line that I highlighted here. The organization has a role, like a clinical role of diabetes educator. They chose to map it to health educator.
(21:04):
They thought, okay, those are similar that send health educator into Cosmos, but we know that 76% of our sites actually picked diabetic education when diabetes educator was in there. So we flag this for this customer and say, your mapping is not what the rest of the community looks like, and we have to do this with every single domain and it’s kind of a pain, but there’s no magic that’s getting through this. We can use AI and help float things up to the list, but we want them to go and fix these things. And the nice thing is that treatment team relationship then is going to be used for clinical care as well, so that as this record can be exchanged, we know that particular mapped element. So that is something that we’re starting to roll out for groups, but this is an implementation of what the OHDSI community tools are doing, the terminologies that we’re using all within this Cosmos portal that our health systems can log into and use.
(21:57):
So I’ll end with a quote here. This idea of data quality is not like we’re going to go on a journey and be done, check the box. It’s like doing something you don’t really want to do, floss your teeth, and you’re just going to have to do it for the rest of your life if you want good data quality. So you’re never going to be done. It’s always going to be happening, but it is very, very important because more and more medical knowledge is going to be based off observational data sets. You’re not going to run a clinical trial for everything. So having high quality data that’s mapped to standard terminology around the world is really critical if we’re going to be making new medical knowledge off these data sets. So that’s my bit.
Carol Macumber, MS, PMP, FAMIA (22:36):
Thank you. So these guys won’t be surprised that I’m going to go a little off script. We have a few canned questions, but as we’re talking, I came up with new ones. So it’s always fun to challenge them on the spot and they’ll forgive me for it later. So Suzy, I’m going to go to you first. Okay. Because you brought up the SNOMED eye chart, you brought up the power of SNOMED for somebody like me, I’m like, this is wonderful. It’s great. It’s a richness. It’s poly hierarchy, meaning there’s multiple parents, you can classify things in different ways. How do you explain to the others who don’t find that so immediately cool, the difference between that being the power of SNOMED versus overly complex.
Suzy Roy (23:16):
So good question. And really again, it’s going to come down to what is it that we ultimately want? We want patient safety, we want better outcomes for patients. So a really good example that just kind of comes to mind. So at the beginning of covid, so we’re talking April, 2020, a large hospital in southeast London, they started seeing that they had patients coming in to the emergency room department and they were displaying what we now know as some of those very clinical signs of that early variant of covid, the cough, the labored breathing. But what they were finding from these emergency room department patients, some were actually having a worse outcome than others. All of their electronic health record in that hospital has SNOMED CT, it has LOINC, it utilizes HL7. And so what they started to do is actually pull out some of those cohort data.
(24:22):
And what they actually found was that patients who were coming into the emergency department who displayed those particular coughs, labored breathing and everything who were of Southeast Asian descent and of African descent. They were finding that they were having the worst outcomes out of everyone. And so what they started to do as those patients were coming in, they were able to triage and they were able to start to provide those particular patients with different care. And then they started to see that those patients were not having as severe outcomes. And really when we start to talk about the use of SNOMED and the power of SNOMED, it’s really that I want people to understand, yes, I geek out. Actually, we were kind of over there geeking out on the poly-hierarchical structure of SNOMED CT and utilizing all of those clinical, the ECL queries. However, while we’ve like that, what we actually really all want is better patient safety, better patient outcomes. And so that’s what I always want people to remember is that yes, SNOMED has a hierarchy. Yes, we harmonize and we now align perfectly with LOINC and other standards. However, what we actually want is that patient safety at the end.
Carol Macumber, MS, PMP, FAMIA (25:42):
Thank you. Yeah. Two for you Dr. Hamlin.
Ben Hamlin, DrPH (25:45):
I just want to say terminology is really cool. I mean if you don’t appreciate that, you’re missing out something, terminology is really cool.
Carol Macumber, MS, PMP, FAMIA (25:54):
All right. Alright, so Ben, collaborative developments and the community aspect of it, I mean does that for you and that community increase trust?
Ben Hamlin, DrPH (26:07):
Yeah. So I love to see actually Phil’s the screenshots so much better than mine. And that’s the context of what we’re talking about. I mean, you’re running structural data analysis on this information from a variety of different data sets. You still require that level of expertise to say, wait, this doesn’t look right because I know what it’s supposed to look like in these summary data aggregate sets in this dashboard, this raises a question. The community is where we can not only help expose that kind of expertise in a more standardized way. So these are the kind of things you see in these sorts of contexts and scenarios. These are the kind of things you should be looking for in these sorts of scenarios. So when they do come up, you don’t need someone with Phil’s level of expertise to say this is not right. This requires more investigation. That’s going to come from two things. One from a lot of what they’re doing in Epic in terms of running these standard structural analyses based in standards and presenting this information in a very user-friendly format. But also in the community side where you have to engage. I disagree on the, we’re not going to do RCTs and everything because the OHDSI community does allow you to do computable RCTs very rapidly on large scale that can produce that kind of information. And as long as we’re capturing that and storing that, and then …
Phil Lindemann (27:26):
On real humans though?
Ben Hamlin, DrPH (27:28):
On real humans.
Phil Lindemann (27:29):
They’re going out and doing randomized clinical trials that are just like pragmatic trials?
Ben Hamlin, DrPH (27:33):
Pragmatic trials at a scale that you could never do in a traditional RCT, right? They’re able to do cross-comparisons, a thousand variables at a time for hundreds of thousands of people. So if you’re looking for
Phil Lindemann (27:47):
But they’re not actually doing this, patients are not consenting into a randomized clinical trial. Yeah, that’s the part I’m saying.
Ben Hamlin, DrPH (27:53):
That’s another conversation
(27:55):
That’s for the patient consent folks to deal with. Alright, that’s the one limiting factor. But the community that is working towards this idea of let’s not only promote best practices and let’s produce the tools to enable these kinds of analyses to happen in a much more reproducible fashion. That expertise is out there. But connecting those people together to do this kind of work under a common cause, I think is where the community is really the strength is of OHDSI community, of the HL7 community, of all the people are the same kind of focus in terms of their same priority goals.
Carol Macumber, MS, PMP, FAMIA (28:30):
Last question, the phrases you use around context of use and limitation of use are very common for folks who’ve ever dabbled in the structured product label and FDA indexing. Is there an effort to, or does it already exist for you to standardize using terminology, the context of use, conditions of use and limitations of use for your phenotypes, for your models?
Ben Hamlin, DrPH (29:00):
So the short answer is yes, but
Carol Macumber, MS, PMP, FAMIA (29:05):
I really did think you were going to say no.
Ben Hamlin, DrPH (29:08):
Well, I’m going to qualify it. It’s a yes, but. I mean again, so the context of use definitions are there, everyone knows what they want to do and how they want to do it. Some of the limitations are in the metadata around those particular context of use and also the reliance on the terminology alone to help them answer those questions. So terminology is very sexy, it is really cool, but it can’t do everything. It requires algorithmic models that have been vetted and tested in real-world data that help associate those particular elements together with the associated context of use metadata and all the other stuff together to say for sure that this is true, this is meaningful. These underlying essences that are the decoding underlying essences within this dataset are really actually meaningful. And they’re not just interesting to geeks.
Suzy Roy (30:06):
I mean there is that. There is that. But also when it comes to especially something with structured product labels, that’s where, and this is what we also see in EHR records and other health data work standards. They’re hard. And while some of us find that kind of stuff sexy, not everyone does. You
Carol Macumber, MS, PMP, FAMIA (30:25):
Guys have said the word sexy
Suzy Roy (30:26):
So many times. I know we need to
Carol Macumber, MS, PMP, FAMIA (30:27):
Stop terminology.
Ben Hamlin, DrPH (30:29):
We’re trying to use words that want to get people excited about this stuff, right?
Carol Macumber, MS, PMP, FAMIA (30:32):
Let’s stop while SNOMED doesn’t do it.
Suzy Roy (30:34):
No SNOMED just does it. No, but really, I mean this is hard. And so that’s some of the challenges in utilizing this. Like yes, if we could all have everything structured and based on an information model, that would be excellent. However, when we start to talk about that, especially in that particular use case with the structured public label, you’re not just talking about SNOMED, you’re also, especially here in the us you’re talking about RxNorm now and then pulling, are we talking about this coming from a particular drug that’s coming from a device like a clicker? And so you’re starting to talk about multiple standards as well. And so while us SDOs really do try to work together, some of these are so specific for a use case, it is hard and that’s going to take time, money not just for the developers of these particular case systems, but also for the standards themselves and also for the implementers. And when you have something that’s not so exciting, there you go, right. That’s hard to try to explain to them and get everyone to buy in. And so there are multiple levels there that really pose some barriers and challenges.
Ben Hamlin, DrPH (31:49):
And it’s a continuous learning cycle. I mean we need Phil to be sharing his knowledge from his instance of application in the standards community because the standards can build a model, but if you don’t take the tires on it and really run it through the wringer with real-world data, it’s okay, but it’s not sexy.
Carol Macumber, MS, PMP, FAMIA (32:08):
Thank you for the segue. So I could get to Phil. I tried. Yeah. Phil, you mentioned the advantage you have from being the content creator, but also being able to utilize it and see it over its lifetime. Pulling on Ben’s topic of trust, how much of that continuum in the degradation of the initial clinical intent do you have to make transparent for the end user and the research community or the analytics people to trust the content when it gets to them?
Phil Lindemann (32:38):
Yeah, I think there’s probably two layers. So one, the more papers that are published on it, the trust builds, but that’s sort of a blind trust. So it’s a little dangerous. One of the things that we’ve said is, so if we can stop the 160-pound baby from being entered in the first place, that’s ideal. Just alert them right in the system. So part of this is us following all the way back to the point of data creation and detecting that is an implausible value. Do not let them put it in there. Then the second step is kind of where we’re focused now, which is let someone more at administrative or a higher level sort of understand where are their problems and then go back. But the final step is we just have to expose this to them. So for example, when you are part of Cosmos, you have local regulations, you might not be able to send HIV results into Cosmos because your government or your local region doesn’t let you do that.
(33:30):
So those will be removed from Cosmos, but the researcher might go in and say, wow, there’s no HIV in this particular area. So we have to be able to expose that to them and say, this may appear like a data quality issue, but just so you know, this group has carved out this particular diagnosis or this particular lab request. So we do need to expose it in front-end tools, which is a much harder thing. But that is our goal, is to be able to take that dashboard and say, okay, I’m in a trend. We were looking at one yesterday, opioid use for migraines in the ED, something you really shouldn’t do. But when we do that study then allow me to say, okay, now apply all the data quality rules, delete all the implausible patient values, delete everyone who doesn’t have a complete record, and then what’s sort of left after you’ve trimmed the fat? That’s the study. So we have work to do. We want to expose it all the way to end users.
Carol Macumber, MS, PMP, FAMIA (34:20):
Okay, so one either or question. You don’t get to say both. So I’ll ask you the same question I asked Dr. Barr yesterday around this concept that it’s essentially for you guys all looking at what you do, it’s an ocean of data and this is a kind of desert of insight. So if you had to choose between more data and better data, which would you pick? We’ll go to Ben first and it’s a short answer. Both. Can’t pick both. No, both more. Why short answer?
Phil Lindemann (34:54):
Because better is too subjective, better means different things to the three of us.
Carol Macumber, MS, PMP, FAMIA (35:00):
Suzy?
Suzy Roy (35:01):
I want more, more. I always want more, but also we are developing better tools so we can start to actually delineate some of those dirty data. So I want more.
Carol Macumber, MS, PMP, FAMIA (35:13):
Phil?
Phil Lindemann (35:15):
I’m going to say better.
Carol Macumber, MS, PMP, FAMIA (35:16):
I was going to say…
Phil Lindemann (35:18):
I want more of it, and I would say the thing that I’m most interested in is capturing more in-depth data elements. I think everybody can say, well, we can look at diagnosis codes and we can look at meds, maybe a little bit of labs, but there is so much specialty nuance data, risk scores, and indexes within a specific disease state that I think are some of the most powerful elements. And oftentimes those are the things most likely to be hiding in a piece of paper that came in at your clinic or it’s a scan getting that information. So it’s not exactly what you’re saying, but that’s what I think is most important.
Ben Hamlin, DrPH (35:56):
I’ll let it slide. Actually, I want to add to that though, because I kind of agree with Phil on one aspect, and I want to use as a scenario that we use a lot in our data quality assessment, which is the pregnant males. We tend to eliminate pregnant males from our certification. We find that as a problem in data, but now we have more specifications to really understand your gender identity could actually classify as male and you could theoretically, I mean you could realistically get pregnant, and so therefore that is no longer, I mean, I don’t know about the 160-pound babies, but pregnant males could be a real thing and that requires more data, but it also requires a better understanding of that more data. So you have to be able to really again, apply the right context.
Carol Macumber, MS, PMP, FAMIA (36:40):
Right. Well, thank you. Give our panelists a round of applause. Thanks for joining us. Thank you.



