Back to IMVC 2026

הסוד ליישור מושלם: למידה עמוקה משנה את פני הראייה הממוחשבת!

גלו כיצד למידה עמוקה מחוללת מהפכה ביישור גיאומטרי של נתונים תלת-ממדיים, סדרות זמן ותמונות. למדו על הגישות החדשניות המאפשרות מודלים קלים, מהירים ויעילים ללא צורך ברגולריזציה מסובכת. פתחו עולם שלם של יישומים מדהימים.

Oren FreifeldOren FreifeldAssociate Professor, Ben-Gurion University of the Negev

פרקים

00:05יישור גיאומטרי: מהפכת הלמידה העמוקה

So I'm going to talk about deep learning for Geometric Alignment, which is a topic that we've been working on in my group for quite some time. And I want to start the talk by giving some examples for Geometric Alignment tasks. For example, תודה. So let's dive in, here is the first example. Suppose you have a pair of 3D Gaussian Splatting Models. More generally, it can be point clouds or other representations in 3D. And what you can see here on the left, can you see the mouth? Yes. So what you can see on the left are two different objects, same category, both are boats, but these are different instances, and they are misaligned in terms of scale, rotation and translation. Here the initialization is just random, and you can see them from two different angles. And the goal is to align them in terms of scale, rotation and translation, what is called a similarity transformation. So this would be an example for such a... for a geometric alignment task. Now, if you are able to do that, then this unlocks all kind of interesting applications, for example, a geometrically consistent object replacement in an image. So here, on the top left, you can see the original car. And suppose we have a 3D model of it, maybe we had several images and we created a 3D Gershons Platting model, and we aligned this model to other instances, then we can replace them and re-render the new objects, and it would look fairly natural. So the red car and the purple cars are both, these are 3D models of real cars. And for the police car, this is actually a synthetic example, but the mathematics is all the same. Now, I talked about an example of alignment in 3D, but we can also do that in the case of one of the signals, known as time series. סדרות אתיות, למדתי שזה נקרא. So, if we have two time series, here in blue and red, then there is a standard solution based on optimization called DTW, Dynamic Time Warping, and it aligns one signal in an optimal way, according to some criterion, to the other. Now, DTW, even though it's very widely used, it has several downsides. For example, it often yields warping paths that are non-smooth. וגם זה לא רובס לנזק ויש לו גם תחומים אחרים כמו זמן הקומפטציה. עכשיו, דבר אחר שיכול לעשות עליו הוא להשתמש בהגדרה הקדושה של לימודים, שמוסדות סלוליות יותר קלות ואתה יכול לעשות את זה גם באינפלנט, אז זה בעצם קל יותר, וזה מה שאתה רואה ברצועה. עכשיו, השניים שראיתי לכם, ב-3D וב-1D, בכל מקרה, דיברנו על העלייה פר-וויזית, אבל לפעמים יש לך קולקציה של נסיגנלים, ואתה רוצה להעליין אותם ביחידה. אתה יכול כמובן לנסות לעשות את זה פר-וויזית, אבל אז הערעורים יקומלו. אז, הצעה טובה היא לעשות את זה ביחידה. ופה למשל, על הראשון ללפת, you see a collection of ECG signals, and you can see they are misaligned in time. As a result, if you want to do some statistical analysis, for example, even computing the mean, what you see here in the solid red line, you will get a smear signal that does not look like an iconic ECG signal, or בעברית אק"ג. And the shaded area represents the variants, and you can see that there is a lot of fictitious variants over here. However, if you are able to take all the signals and align them in time, then once you average them, you can see that you get something which looks more familiar, like the iconic ECG signal. Now, it is a machine vision conference, so I'm bound to show some images, and you can also talk about alignment of images. For example, you want to create a panoramic image, you have two images, and this is not from our work, I think this is from an open CV tutorial, and you can take two images, find some corresponding points between the images, and then estimate the transformation, such as a homography, and apply them to create a panoramic image. And here, too, in the case of images, just like you would like to align pairs, sometimes you would like to align a collection of images, או שזה נקרא "קונג'ילינג" בכל מיני מקומות. אז כאן, בין כל קטגוריה, למשל פלנות, גלעדס או בייקס, אפשר לראות שכל גלעדס מיוחדים, והם מיוחדים, ואז אנחנו משקיעים סלולים רחוקים כל יום כל האימג'ים בקטגוריה הזאת. עכשיו, ההפיכה שאני רוצה להציע פה, יש שתי קבוצות ראשונות. הראשונה היא שאתה צריך לעשות דברים, אני אגיד, שאתה צריך לעשות דברים בצורה שגיאומטרית יודעת, והשנייה, וזה אולי יותר קונטרוברסי, is that whenever possible, you should do this in a way which is regularization free, and I will explain what I mean by that. Now, just to give some motivation behind this approach, in this example I just showed you, with the bikes and the joint alignment of collection of images, previous methods, and these are fairly recent, from three years ago, So previous methods used to take more than an hour to jointly align a collection of 30 images. And the two methods in the bottom are from our group, and these are much faster and more lightweight. So we have... far fewer parameters, it requires less epochs, and you can see the running time. But what I want to draw your attention to is this column. The number of... The number of hyperparameters. And because our approaches are based on regularization-free, regularization-free, formulations, then we can do this in a way... שאין להם פרמטרים חדשים, לפחות בקונטקסט של רגולציה. וכתוצאה מזה, אין לנו צורך להצטוות את הפרמטרים האלה על ידי סט-דאטה. עכשיו, נראה שמה שראינו בטלפון בעצם עוסק בגללו. ההצגה נותנת מודלים שהם קצת קטנים, אפישיים, בעיקר עם תוצאות טובות, והם יותר אפשריים להתאימו ולא קצת חזקים. כשאתה הולך מאחורי דאטה סט להשני דבר, הם לא מגיעים כל כך יחידות. כמו שאמרתי, אנחנו עובדים על כאלה בעיות כזה לפני זמן מסוים, בכל מיני קונפיגרציות ומסגרים ובדיקות דאטה רבות, כמו סטיימס סיריוסט, או וידאו, או וידאו מסוימים, אימג'ס, או 3D, פוינט קלאודס או 3DGS. And we considered all kind of transformations, such as non-linear time warping, F-line transformations, HOMOGRAפיז, similarity transformations, or non-regit transformations. And as I mentioned, this is relevant in 1D, 2D, 3D, and so on. So, what I'm going to do now, I'm going to touch upon two key components that we use. The first one we did not invent, this is from DeepMind, and this is called STN. And STN is a differentiable module that learns input dependent spatial transformation for a downstream task, and then it applies it to the input. And the second component is what we call CPub, continuous piecewise F-line based transformations. And I don't have time to go into all the details, but in a nutshell, here how it goes. There is a mathematical concept called difumorphism. A difumorphism is a differentiable invertible map with a differentiable inverse. Now, CPAP is a particular case of difumorphisms, and it's a finite-dimensional family, it's a finite-dimensional parametric family of efficient and highly expressive difumorphisms. In particular, you can do the evaluation of both the transformation and the gradient with respect to the parameters. You can do this very fast, and in fact, in 1D we even have closed form. So in this particular example, in 1D, we can use this kind of machinery to parameterize the non-linear time warping that took us from this signal to the one below. עכשיו, אם ניקח את שני הקונספטים האלה יחד, ה-STN וה-CPAP דיפיומורפיזם, ונעשה פורמולציה בלי רגולריזציה שאני אספר לכם עליהם במהלך דבר, אז נוכל להביא את הסיסטם המסגרת הזה, שהוא מאוד קל-חזק, והוא מבחינת את הפרנספורמות של CPAP, ולאחר מכן מבחינת אותן לתוצאות, וכתוצאות כך נוכל להגיע לסיגנל. אז זו הפיגורה שראינו לפני כן, ללפת, And it's the same network that works on both, on several classes, in this case, two classes. So during training, we know class labels, but we don't know the ground truth alignments. And during test time, we don't know the class labels, of course. So it's the same network that jointly aligns each class separately during test time. עכשיו, אם אפשר לעשות את זה, אפשר גם להפוך את האידאות האלה לוידאו, ובשמחה אהיה יכול להכין את הוידאו שלי, שזה מאוד מגניב. אוקיי. I don't know anything about baseball, so I will go with a pitch. So, what you are going to see here, in the top, these are the original videos, they are not synchronized, and what you see at the bottom, we synchronized all the videos, even though these are different players, different viewing angles, different cameras and so on, we can still synchronize them. Let me just run it one more time. So at the bottom, this is after synchronization, and at the top, this is before synchronization. from moving cameras and synchronize them in space, I mean, jointly align them in space, to create a panoramic moving camera video. So, having worked on this for quite some time, I want to share some geometric insights, what makes this thing actually work. When you consider transformations, first of all, invertibility is actually key, because it offers several benefits. First of all, it leads to much more stable optimization. So when you train an STN, if the transformations are constrained to be invertible, you can have a much higher learning rate and... ‫האופטימיזציה היא פחות גדולה. ‫בשלב שני, יש פה קונקציה ‫למה שקוראים לסטרוקציה ליג-גרופית, ‫שבאה עם אופן קל ‫לפרמטריזם את המסגרת, ‫בשלב שזה מסגרת שלא ליניארית. ‫ואנחנו נצטרך את האינברטיבליות ‫של פורמולציה בלי רגליזציה ‫שאני אציג תיכף. ‫עוד פעמים, זה הרבה יותר אפשר ‫למצוא פרנספורמציה קטנה ‫מפרנספורמציות גדולות. אז אם אתה צריך תהליך רב זה יותר טוב לקמפוז כמה תהליכים קטנים. ויש עוד אישור דליקטי עם פליפים, כמו רפלקציות. בעצם מה שאומרים זה "ברוט פורס", אבל בעצם יש דרכים לעבור על זה, ו"ברוט פורס" אינו הדרך. אוקיי, אז טיפולית מה שיש בקבוצות העלייה הקואליציונית יש לך איזשהם פונקציות חוסרות שיש לך קולקציה של אימג'ים, כאן II And you want to minimize some discrepancy measure between them and some latent atlas, which is also called the template or the mean image. Now, the problem is that it's very easy to get a poor global optimum for this. Usually, we are used to think of poor local optima, but actually, poor global optima are also a problem. So, in this example, if you want to jointly align all the images and you want to minimize the variance, למשל בכל פיקסל, אז אתה יכול להשתמש באימג'ים בפקט, ואז הוויינס יהיה זירו, ואתה משקר, אתה קיבל פונקציה של זירו. אבל, כמובן, זו לא סיבה חיובה. אז מה שבאנשים עובדים לעשות זה תרגולת אדר, אז יש לך פרמטר היפה למדא, ואתה צריך לשנות אותו. אבל, בקונטקסט הזה, אלימות קבוצות, that such regularization is actually a bad idea. First of all, it's not even needed. Here is an example of one of our methods, which is regularization free. And second is that people are often unaware or tend to underestimate how problematic the use of such terms is. The first problem is that you need to tweak the hyperparameters per data set and sometimes even per class. The second problem is that it just makes the valuation of the loss more expensive. and optimization harder than it should be. And this limits the practicality of using these methods. Moreover, these difficulties drove people to use expensive, high-capacity backbones, as well as memory-consuming, high-dimensional, deep features such as Dino. And again, this is not actually needed. Lightweight models can do just fine. And I'm not against using Dino here or in general, but I just argue that there is a better way to formulate a task regardless if you use features like Dino or not. So, we developed several regularization-free methods, and I will only show one of them here. And these are based on new loss functions, and these losses are invariant to a global transformation. This invariance simplifies the optimization while maintaining a unique solution in terms of the correspondences across the N images. And once the optimization is done, we can choose any arbitrary global transformation that we want for the purposes of visualization. And this does not affect the quality of the solution. So here is an example, and המסגרת המשמעותית כאן היא שאתה רוצה שהפונקציה שלך תתמודד לסביר את הדעת. כאן, במקרה של לנסות למנוע איזשהו פרסומת, אנחנו מנסים פרסומת של איניברס, זאת אומרת, אנחנו לוקחים שתי אימג'ות, ii ו-ij, ונפנה את הפרסומת של איזו אימג'ה, ואז את הפרסומת האיניברס של האחר, ואז אנחנו נבדוק אותן. עכשיו, הפורמולציה הזאת מפרסמת אותנו להסביר את כל ההבנות, וכתוצאה מזה אנחנו לא יכולים לקבל סלולים טריוויאליים כמו אלה שראינו קודם, עם הפור גלובל אופטימום. וזה נראה כך, אפילו שזו פורמולציה כל כך ספציפית, זה בעצם מה שהגיש לנו את הסלול שראיתי לך קודם. בלי רגולציה, זה פשוט עובד, והשטח הוא הרבה יותר קטן, וכל הבנות שהצגתי על הדוכן לפני כן. In this particular example, we do use Dino, and we use a frozen Dino encoder and then some dimensionality reduction using auto-encoder, and then we apply the STN over here and refine the features. And the features that we learn know something about alignment. So in the middle row, you see the Dino features, using PCA for visualization to dimension-free. ומה שאתם רואים פה זה הפיצויים החדשים אחרי העלייה, ואתם יכולים לראות, למשל, שהאירות הן כולן רגעות, אז זה הרבה יותר אפשר להעליין מאשר בקריאה השנייה. ובפיגור שראיתי לפני כן היה את הקומפוננט הזה, וזה הוא תוכנן STN, וזה קשור למה שדיברתי עליו, שאנחנו רוצים שהתרנספורמציה תהיה קטנה בגלל שזה הרבה יותר אפשרי. We apply several recurrences of the STN and each time we predict small transformation and then just compose them. And finally, SpaceGem, the model I just showed you, uses Dino and Dino are dense feature maps, so it cost you, especially in terms of memory, and it turns out you don't even need that. And this is a more recent work that was presented today by Omri earlier, Omri Hirsh. And here we replaced Dino, and instead of using dense feature maps, we used spars representation and graph neural network, but the motivation is very similar, we still use an inverse compositional loss, and this is what lets us go for minutes to seconds. So, to summarize, Geometry matters when you solve geometrical alignment problems. Instead of just throwing very expensive deep backbones and expensive features, you should pay more attention to geometry, and whenever possible, it's better to use regularization-free losses. Thank you.

00:20קסם התלת-ממד: יישור מודלי Gaussian Splatting
01:26יישומים מפתיעים: החלפת אובייקטים חלקה בתמונות
02:15מיישור סדרות זמן ל-DTW: פתרונות חדשניים
03:22מעבר ליישור זוגות: יישור קולקציות שלמות
04:29יצירת פנורמות מושלמות: יישור תמונות מתקדם
05:40שני עקרונות מנחים: גיאומטריה ואי-רגולריזציה
06:04למה לוותר על רגולריזציה? מהירות ויעילות
07:20מודלים קלים וחזקים: היתרונות של הגישה שלנו
08:23המרכיבים הסודיים: STN ו-CPAP
09:40מסגרת פורצת דרך: STN + CPAP ללא רגולריזציה
10:36סנכרון וידאו בזמן אמת: המהפכה הבאה
12:32תובנות גיאומטריות: חשיבות ההפיכות
13:51האתגרים ביישור קבוצות: אופטימום גלובלי והרגולריזציה
15:51הפסד ללא רגולריזציה: פתרונות חכמים
17:23מעבר ל-Dino: יישור יעיל עם תכונות דחוסות
18:22העתיד כבר כאן: ייצוגים דלילים ורשתות גרפים
19:00סיכום: גיאומטריה היא המפתח ליישור מוצלח

יישור גיאומטרי הוא תהליך שבו מתאימים אובייקטים או נתונים שאינם מיושרים מבחינה גיאומטרית. המטרה היא ליישר אותם במונחים של קנה מידה, סיבוב והזזה, מה שמכונה טרנספורמציית דמיון. לדוגמה, ניתן ליישר שני מודלים של ספלאטינג גאוסי תלת-ממדיים או ענני נקודות.

יישור גיאומטרי פותח מגוון יישומים מעניינים. לדוגמה, הוא מאפשר החלפת אובייקטים עקבית גיאומטרית בתמונה, כמו החלפת מכונית במודל תלת-ממדי אחר. בתחום האותות, הוא מאפשר ניתוח סטטיסטי מדויק יותר של סדרות עתיות, כמו קבלת אות אק"ג מוכר לאחר יישור. כמו כן, ניתן להשתמש בו ליצירת תמונות פנורמיות מרובות תמונות או לסנכרון סרטונים ממצלמות שונות.

הגישה מבוססת על שני עקרונות מרכזיים: ראשית, יש לבצע את הפעולות באופן מודע גיאומטרית, תוך מתן תשומת לב רבה לגיאומטריה של הבעיה. שנית, ככל האפשר, יש להשתמש בפורמולציות ללא רגולריזציה (regularization-free). עקרונות אלו מאפשרים פתרונות קלים, יעילים ועם תוצאות טובות יותר.

גישה ללא רגולריזציה מועילה מכיוון שהיא מבטלת את הצורך בכוונון פרמטרי היפר-פרמטרים לכל סט נתונים או קטגוריה. היא מפשטת את האופטימיזציה בכך שהיא הופכת את הערכת פונקציית ההפסד לפחות יקרה ומקלה על תהליך האופטימיזציה. בנוסף, היא מאפשרת שימוש במודלים קלים ואינה דורשת שימוש ב-backbones יקרים או פיצ'רים עתירי זיכרון.

המסגרת משתמשת בשני רכיבי מפתח: הראשון הוא STN (Spatial Transformer Network), מודול דיפרנציאבילי שלומד טרנספורמציה מרחבית התלויה בקלט ומחיל אותה. הרכיב השני הוא CPAP (Continuous Piecewise Affine based transformations), שהיא משפחה פרמטרית סופית-ממדית של דיפיומורפיזמים יעילים ומאוד אקספרסיביים. שילוב STN ו-CPAP, יחד עם פורמולציה ללא רגולריזציה, יוצר מערכת חזקה וקלה.

היפוך טרנספורמציות הוא מפתח מכיוון שהוא מציע מספר יתרונות. ראשית, הוא מוביל לאופטימיזציה יציבה בהרבה, ומאפשר קצב למידה גבוה יותר באימון STN. שנית, יש לו קשר לתיאוריית חבורות לי, המספקת דרך קלה לפרמטריזציה של טרנספורמציות לא ליניאריות. בנוסף, קל יותר למצוא טרנספורמציות קטנות מאשר גדולות, ולכן עדיף להרכיב מספר טרנספורמציות קטנות.

הגישה שלנו משפרת משמעותית את היעילות והמהירות. שיטות קודמות נהגו לקחת יותר משעה כדי ליישר במשותף אוסף של 30 תמונות, בעוד שהשיטות שלנו מהירות וקלות בהרבה. הן דורשות פחות פרמטרים ופחות תקופות אימון, ומאפשרות לעבור מזמני ריצה של דקות לשניות.