Back to IMVC 2026

האם צילום רנטגן פשוט יחליף סריקת CT לאבחון PE?

תסחיף ריאתי הוא מצב מסכן חיים הדורש אבחון מהיר, אך סריקות CT יקרות ואינן תמיד נגישות. גלו כיצד מודלי AI גנרטיביים מצליחים לשחזר סריקות CT תלת-ממדיות מצילומי רנטגן דו-ממדיים פשוטים. פריצת דרך זו מאפשרת זיהוי מוקדם ומדויק יותר, ומפחיתה באופן דרמטי את הצורך בבדיקות CT מיותרות.

Noa CahanNoa CahanComputer Vision Researcher, Tel Aviv University

פרקים

00:05הסכנה הנסתרת: מהו תסחיף ריאתי ולמה הוא קטלני?

Hi, I'm Noah and I'll present our works on generative models to enhance pulmonary embolism or pe. We have a few collaborators for these works. The first one is Professor Giriz from Tel Aviv University and two medical centers, Shiva and Sorotsky. So, in our research we would like to augment the classical classification or detection of pulmonary embolism or pe. PE is a life threatening condition and is very time sensitive. We use a special CT called ctpa. It's a CT with contrast agent to detect this pathology. And what you need to know about it is it can be seen in the CTPAs, but it cannot be seen at all from chest X rays. In the images here you can see examples of PE clots where the contrast agent dyes the pulmonary arteries in a bright white color and the PE clots appear as gray blockage areas. In our research we address a key issue when in the PE diagnosis process, which usually starts when a patient arrives to the emergency department with some kind of lung related issue. After triage, the patient is then sent to multiple testing, including a chest X ray or cxr. And then after testing and further testing, a decision needs to be made whether to send the patient home or keep him in the hospital for further testing, including a CTPA if PE is suspected. When we talk about imaging the lung area, we have mainly two imaging modalities, X rays and CTs. X rays are 2D imaging. They're very fast accessible, they're cheap, but they're 2D imaging, which allows us to have only a limited visibility. Whereas in cts we have a much, much better understanding of what's going on anatomically. But CTs have a much higher amount of ionized radiation to the patient. It's also very expensive and they're not always accessible. If we add a contrast agent, we can see the deficits in the lack fillings which are diagnosis for PE. So basically we would like to use the CXRs as much as possible, but it's not sufficient for all cases. Specifically for PE diagnosis, X rays are not enough. So in this research we aim for CTPA level insight, but with CXR level accessibility to enable a broader and much earlier PE triage. So for this research, we gathered patients who underwent both chest x rays and CTPAs. We then want to take only a single 2D chest x ray and use some kind of generative model to reconstruct and generate the full 3D CTPA scan and then use these generated scans to enhance the PE classification. Prior work includes either automated detection of the PE clots themselves, but from CTPAs, not from CXRs or some works who did the generation, but from DRR to ct. DRR is a kind of artificial X ray that is reconstructed from the city itself. Okay, so for our generative model we chose very simple latent diffusion model or ldm. I'm sure all familiar with this. I'll go over this model a little bit and show the differences between this model and our model. We chose this model because we hypothesized that the X ray to CTP generation acts more like a text to image rather than image to image translation. So in an LDM, at first we input a 2D image into the VAE to get the latent representation which is then noised up and input to the classical 2D unit. And we use the encoded text by usually clip or something like that for conditioning our results to get our matching generation. So in our model we made a few changes. At first our 2D image is now a 3D volume which is input to the same VAE slice by slice to create our latent representation. Our latent representation is now not 2D latent but a 3D one. So the whole thing is expanded from 2D to 3D. Also our unit is now 2D is 3D instead of 2D. And as before, we use a conditioning instead of the text, we input the embedded encoded CXR as conditioning. It's same as before, concatenated through the unit's layers to get our wanted generation result. So during training we input our CTPA and our cxr. But in inference in the red arrows, we only input our noise and the CXR to get our matching ctpa. That is, as we request, we also add a discriminator to enhance appearance and APA classifier to enhance our classification results. The full loss can be seen here below. Our data set includes around 900 patients who underwent both the CTPA and X ray and the pes. We extracted only a single binary label from the whole 3D volume. From the radiologist report indicating whether the patient had PE or not, I wanted to show some results. So on the left here you can see our CXR input and in the middle you can see the ground truth of the full CT scan. And this is our generated result. As you can see, it closely follows the ground truth, but there are small variations between them. Here are more examples from our test set. I thought it would be interesting to look at the bottom right example here where a patient is presented with a condition called pleural effusion where a bright white color is seen in the chest X ray. This is a condition where fluid fills up the lungs. This is not a pathology our model was trained on. And yet we see that the model is able to generate it at the correct location. Again, small variations such as texture, for example, for. But overall, it's interesting to see this result. Okay, so what about PE clots? Is our model able to really generate the PE clots which are very small? So the answer is overall, yes. You can see samples here and their ground truth. Slide it and you can see the PEs are indeed generated at the correct locations. For explainability, we present here t sne results of 2D and 3D t sne of our generated samples. In blue, this is our ground truth. In orange, you can see the generated 3D CTP samples in a conditional setting. And in green you can see generated CTPA samples, but in an unconditional setting, meaning we didn't feed any conditioning. So just generate CTPA samples without any conditioning and you can see that the blue and orange dots closely follow each other. Okay, so for our PE classification results, let's first look at their baselines. So for CTPA only classification, we achieved 90% NIUC. For CXR only classification, we achieve 69%. But if we take our CXR generate from it the full 3D CTPA, we enhance our result to 80% in AUC. So overall we improve our results by 11% in AUC. But more importantly, we have a higher number of true negatives, which basically means we can reduce the number of unnecessary CTPAs performed. Okay, so in this work we worked really hard for the generation, but actually our goal was classification. And this is a little bit of an overhead. So in a successive network and work, we wanted to further enhance our PE classification results, but this time without the need for image generation. So first we trained unimodality encoders for each modality separately and get a 1D representation of each modality. We then trained a very lightweight 1D diffusion prior to generate the CTP embeddings from the CX R1s and then again use these embeddings to enhance our classification results. And we further enhanced our results by a total of 13% in AUC. These are state of the art results for PE classifications from cxr, which was never done before. And we achieved this with an orderless of trainable parameters. So these models are the first CXR to CTPI diffusion models. They show potential for earlier identification for PE and cxr. And these Works and models can be generalized to any other two domains and can be generalizable. And this can be used as a first step for AI for clinical decision support. We plan to scale up to multi center cohorts and expand to other medical modalities. Our lab, currently we have a lot of openings in our lab. That's it. Thank you.

00:26אתגר האבחון: המסע המורכב של חולה עם חשד ל-PE
01:58רנטגן מול CT: איזו טכנולוגיה עדיפה לאבחון ריאות?
02:53החזון: רמת דיוק של CT עם נגישות של רנטגן
03:19המודל הגנרטיבי: איך בונים תלת-ממד מדו-ממד?
04:07מאחורי הקלעים: מודל הדיפוזיה הנסתרת (LDM) בפעולה
05:07הטוויסט שלנו: התאמת המודל הגנרטיבי לצרכים רפואיים
05:58איך המודל לומד? תהליך האימון וההסקה
06:32התוצאות הראשוניות: האם הצלחנו לשחזר CT מורכב?
07:55האם ה-AI יכול לראות קרישי דם זעירים?
08:21הבנת המודל: איך ניתן להסביר את ההצלחה?
09:07פריצת דרך באבחון: שיפור דרמטי בדיוק ה-PE
09:52השלב הבא: אבחון משופר ללא צורך ביצירת תמונה מלאה
11:00עתיד הרפואה: AI ככלי תומך החלטה קלינית
11:43שאלות ותשובות: האם יש מידע נסתר בצילום רנטגן?

Before you started, what was your reason to be confident that you were able to X ray images? If what do you see there that is translatable into the generated? So the question was what made me believe that this model is even possible to create? And if there is anything in the CXRS that is indicative that can be used for translation or something like that, I would even tell you how to have something there, right? Yeah. What is the something and how did you find it? Okay, so if you ask a radiologist if they can detect pes in the CXRs and we even did like a user study for this, so the doctors get annoyed because you cannot see anything. It is not done. It's not something that is used in practical like every day to day work. But you saw that the baseline was that we do achieve 69% in AUC when we try to classify Pes from the X rays. So it's not nothing. There is something there. Right? It's not a 50, you don't get like a 0.5 AUC. So there is probably. So this is like my intuitive answer to you. There is some kind of information from the X rays. We see it, but it is very hard to extract it or enhance it or be used somehow. And we believe that the Latin space of the generative model and in general generation for a pretext task really gives us a very meaningful Latin space representation that allows us to enhance the classification results. We can go deeper later on this. Okay.

13:58שאלות ותשובות: האם המודל מדמיין? אתגר ההזיות

Hi. Since you have ground truth and you're generating stuff, so you're generating 3D with some certain degree of hallucination. Did you try decoupling like the intrinsic ambiguous of the classifier and the hallucination? You mean if I to measure the classification error related to ambiguity and the one related to hallucinations. Okay, so the question was. Actually, there's a microphone. Sorry. Okay, so no, we didn't do this. We didn't quantify the hallucinations. We do know that there are hallucinations. You have to remember this is the first work of its kind. So we kind of just show potential. We didn't quantify the hallucinations, but it could be very interesting. I'm guessing it could enhance our classification results even further. Just to get a ballpark estimate, how many images did you use for tuning the conditioning the model? Basically, because it's not single. Seems very easy to find pairs of the same patient with X ray and ct. Yeah, exactly. So our data set is actually, there's no public data set that has paired images from chest X rays and CTs. We only have 900 pairs of patients. But we did do a massive pre training with just a CTPA data set.

המחקר עוסק בשימוש במודלים גנרטיביים כדי לשפר את הסיווג והאיתור של תסחיף ריאתי (PE). המטרה היא להשיג תובנה ברמת CTPA עם נגישות ברמת צילום חזה (CXR) כדי לאפשר מיון מוקדם ורחב יותר של PE.

תסחיף ריאתי (PE) הוא מצב מסכן חיים ורגיש מאוד לזמן. בעוד ש-CTPA מציע הבנה אנטומית טובה, הוא כרוך בקרינה גבוהה, יקר ולא תמיד נגיש. צילומי חזה (X-rays) זולים ונגישים אך מספקים ראות מוגבלת ואינם מספיקים לאבחון PE.

המודל מקבל צילום חזה דו-ממדי יחיד ומשתמש במודל גנרטיבי, ספציפית מודל דיפוזיה לטנטי משופר, כדי לשחזר ולייצר סריקת CTPA תלת-ממדית מלאה. סריקה תלת-ממדית שנוצרה זו משמשת לאחר מכן לשיפור סיווג תסחיף ריאתי. צילום החזה משמש כקלט מותנה, בדומה לטקסט במודל טקסט-לתמונה.

שימוש במודל הגנרטיבי ליצירת CTPA תלת-ממדי מצילום חזה שיפר את תוצאות סיווג ה-PE מ-69% AUC (צילום חזה בלבד) ל-80% AUC. זהו שיפור של 11%. חשוב מכך, הוא הוביל למספר גבוה יותר של תוצאות שליליות אמיתיות, מה שמפחית את הצורך בביצוע סריקות CTPA מיותרות.

כן, המודל מסוגל לייצר קרישי תסחיף ריאתי (PE), למרות שהם קטנים מאוד. דוגמאות מראות שה-PEs שנוצרו מופיעים במיקומים הנכונים בהשוואה לאמת הקרקע (ground truth).

כן, עבודה עוקבת שיפרה את סיווג ה-PE ללא צורך ביצירת תמונה מלאה. היא כללה אימון מקודדי יחידניות (unimodality encoders) לקבלת ייצוגים חד-ממדיים לכל מודליות, ולאחר מכן שימוש ב-prior דיפוזיה חד-ממדי קל משקל ליצירת הטמעות CTPA מהטמעות CXR. גישה זו שיפרה עוד יותר את התוצאות ב-13% ב-AUC, והשיגה תוצאות חדישות לסיווג PE מצילום חזה.

למרות שרדיולוגים בדרך כלל אינם יכולים לזהות PE בצילומי חזה, סיווג בסיסי מצילומי רנטגן בלבד השיג 69% AUC, מה שמעיד על קיומו של מידע בסיסי כלשהו. החוקרים האמינו שהמרחב הלטנטי של מודל גנרטיבי, המשמש למשימת הקדמה (pretext task) כמו יצירת תמונה, יכול לחלץ ולייצג מידע עדין זה באופן משמעותי כדי לשפר את תוצאות הסיווג.