Back to IMVC 2026

לפענח את המציאות: סודות ראיית המחשב והבינה המלאכותית

ראיית מחשב עברה כברת דרך, אך האתגרים נותרו מורכבים. גלו כיצד אנו מתמודדים עם נתונים לא מסומנים, מציאות פרקטלית וקבלת החלטות בבינה מלאכותית.

Yonatan WexlerYonatan WexlerCTO, Q

פרקים

00:05ראיית מחשב: מסע אל העבר והעתיד

I'm going to give a nice overview presentation about the field of computer vision. And it started many, many years ago as a project to get computers out in the world. And nowadays, I guess for those of you from this millennium, you kind of forget about it. And so I thought I'll do a quick review and recap of how the field evolved and where is it going to, presumably with this. First, for those who didn't hear, we had a company that's unknown now part of Apple. I wasn't sure what we're doing, so I looked at the web, that's what they say. And now the field of machine learning, which kind of swallowed the field of computer vision, has this figure and I bet most of you have seen it, all of you. So we have reds which are bad, and blue which are good, and we're trying to find the boundary between them. And it looks like this, right, in all the textbooks, but reality is more complex, right? Sometimes we can imagine it like this, right? So it's, you know, there's some outliers and some things don't really fit, but it doesn't really capture everything. And everyone here who's dealing with something with computer vision and those, you know, some of the presentations you saw earlier, it doesn't really work like that because the big question is how do we represent the world? What is the world? Right? And this is a big question that was, you know, bothering people for many, many years and will surely continue to bother people. One of the early debates was, is the world discrete or continuous? So in the beginning, the early days of computer science, there big debate about it. There were the Alan Turings of the world that said the world is discrete because it's much easier. A 0 is 0 and 1 is 1. But there were many opponents that said, no, but reality is continuous, it cannot be discrete, it has to be analog. We know this is the world we live in, but it's much harder to represent it and build devices on it. One of the interesting twists is that after Alan Turing won and all of you learned about Turing machines, there was a guy called Michael Rabin that some of you may have been lucky to learn from him. He was a professor in Hebrew University and he passed away this year, was one of my best teachers, and he showed that a probabilistic Turing machine can do stuff that we do now every day. And one of them is find prime numbers, which is the, the core of all of digital communication. So probability is very, very important. So probability is a way to make the world less discrete. And more continuous, but also probabilistic. There are harder questions, right? So the decision boundary between the good and the bad, we know it is not aligned. We know that. And for those of you who don't know this picture, it's the Mandelbrot set. It's a very simple function. You take a number, you square it, and you add a constant and you see how long it takes before it starts going to infinity. And you count these iterations. And the nice thing about the mandible set or any fractal function is that it doesn't matter how much you zoom in, you will always find two points that behave the same, and between them, the whole straight line between them has a whole world that behaves differently. And it sounds like a stupid, silly question or definition or nuance, but this also happens in the real world. For example, here's a story. So there's this kid who broke the window of a store, right? So it must be a thief, right? But it turns out that he was running away from a guy that was trying to kill him. So maybe the kid is a hero for trying to managing to run away. But it turns out that the kid was actually stealing from the store and the store owner was trying to catch him. So maybe the kid is a villain, but it turns out that the store owner was stealing from the kid's family for all these years. And it can go on and on and on and on, ad infinitum. So that, you know, reality is fractal as much as our fractals, right? Yeah. Here's another one. Another one. So clever machine learning. People say, no problem, we need more data, right? So there's a big deal about getting more data. And maybe if we have enough data, we can capture the complexity of the world and then we do whatever we can do with it, right? And we all know the games. But then big data has a big problem, which is labels, right? So if you have billion samples, you want to label them somehow. And you know, you can use a lot of tricks. One of them is unsupervised learning. One of them is to say, okay, I know something about the data. I'll use clustering. I learned a lot about pathology. I'll use my experience and I'll make something out of that data. Or I'll take millions of doctors and they'll tell me what the answer is and I can learn from it. And this can work to an extent. But getting labels is not easy. I'll show you two quick examples. If I asked you, do you See squares or parallelograms? What do you see? Squares, 10 squares. Parallelograms, same 10. Okay, thank you. So it's not trivial, right? You ask people, I don't know. Here's another one. It's my first paper from the previous millennium. So we took two pictures of an object and we tried to find a square function that hyperbola that can approximate the 3D. Okay? So if you look at these pictures, you say, okay, I see a statue, a famous statue from two different angles. If I now apply the two transformation, most people will say, well, I see someone dancing, right? So the label here is not trivial, right? So it's just two simple cute examples. Why Labeling is not easy, but it can get much harder than that, right? So imagine you have unlabeled data. And it's not only unlabeled, it's very high dimensional. And it's not only high dimensional, it's a time series, right? So it's not one decision, right? You have a whole history of every sample that is also relevant to an extent that you may or may not know. And it's noisy. And then what do you do? Right? You do unsupervised learning on this. You're going to basically model the noise. So here's a visualization how real world data actually looks like. There's tons and tons of signal and it doesn't matter what you do. It can be medical, it can be self driving, it can be anything. But sometimes you get a feeling that there's some data that's hidden there, right? There's some signal that's hiding there. You get that hunch. But what do you do? What can you do to get the actual data out of there, right? So again, you can say, I know something about this data, I'll make the real thing float out and surface. Or maybe I can do some eigenvalue decomposition where I prefer some eigenvectors over the others, but it's not trivial. So I'll show you a quick solution that we of a paper we published two years ago, and you're welcome to read it later, but I'll just give you a quick taste of it because you are brave and it is the end of the day. So let's think about this very simple subtask. Okay? So let's say I have tons of examples, a thousand. And let's say I know that they come from, let's say 10 classes of some sort. And someone gave me two classifiers that's supposed to be helpful. One is F1. One is F2. And all I want to know if should I use both of them or not? Okay, It's a very simple question, right? How can I do it? Right? And how can I do it in a way that, you know, will not take me time? So I can do the following. So I assume for now the two classifiers are metric learners which are equivalent. I take all the points and I map them with F1 and map them with F2. They both map into whatever spaces they have. And then I'll take a sample from that thousand examples and I say, well, what do I think about this sample? Let's say it's class A, and if I don't know, I'll just throw it back and I take another one, and I'll take another sample and I see that it's class B. And I take another sample, I see that it's class C. Okay? So I basically had to label three cases that actually know what they are. I map those cases to the relevant spaces that F1 and F2 induce, and then I can say, you know what? These are as good as it gets, right? So I'll ask F1, who is the closest to that A I labeled and I label all of these guys as A and B and C. And then I can map them back to the original space. And now for every sample, I have two opinions. One classifier says A, B or C and the other would say A, B and C. And now what I can do is I can build a histogram, a two dimensional histogram if it's two classifiers and I can look at it. So here's one case. So F1 says everyone is B and F2 says everyone is A. Right? And what can we tell about that? They're both useless. Okay? So it's obviously doesn't matter if I use them or not, if I combine them or not. I can also have another case where I check the histograms and I see that they are very, very correlated. They don't add information on top of each other, right? And this is, and I only labeled here 5 classes out of the 10. But I can already say they're useless. I could have something like this where F1 says that everything is almost everything is B and F2 says almost everything is A. But I can see that they have mutual information compared to each other. And this is a discrete formulation on a very problematic case. And we can talk about it, but I won't say it now, we can talk about it later. And now the question is, can I tell something about this, right? So imagine a very big matrix which is mostly zeros, because that's how they are. And I want to know if the two classifiers are mutually informative without having labels, right? In this case, I only took five samples and gave them some class out of whatever data set I have. So what we did is we formulated it as a bipartite graph and we built a very efficient algorithm that can very quickly assign possible classes to this histogram. And again, I won't go into details, but I'll just show you that we tested it on, on a very big face recognition data set because it's a known and available problem. We chose 100 random people and we labeled them out of 100,000. So it's. Sorry, out of 10,000. So it's 1% of the labels, right? It's much, much less than 1% of the data. Just 100 samples labeled. And we see that the method can. Oh, well, the method is very correlated with the success that we could. Right? So this is labeled data set. So it's easy to measure. But we took unlabeled data for the method and we take two random classifiers and we can tell if it's going to be beneficial to combine them or not. So this is just one bit of unlabeled data that you can use, but you can think how to extend it to anything else. So to conclude this method, and to conclude this day, this is a new way and there's many more ways to enhance what we can do. And if you think about the previous presentations that you had and you saw, they all fall into this. So you have a very complex world. You can use domain knowledge to extract the value out of it, or you can say, you know what, there's some things that I can say about the results that I'm going to get, because the world is continuous, but the decision is discrete, Right? It's either class A or class B, but the representation itself is continuous. And this opens up a lot of. A lot of domains that can benefit and a lot of jumps that we can make in the AI and in the world. And we're going to use it. And. Yeah, and thank you. So if there's no questions, I'll tell you that we're going to build a very big AI team. And if you guys are, you know, want to contribute, you're all welcome to contact us. Thank you.

00:51כשהלמידת מכונה בלעה את ראיית המחשב: האתגר הראשון
01:34איך לייצג את העולם? דיסקרטי מול רציף
02:17מייקל ראבין: ההפתעה ההסתברותית ששינתה הכל
03:00הגבולות הבלתי צפויים: כוחם של פרקטלים במציאות
03:48הילד ששבר את החלון: מציאות פרקטלית בחיי היומיום
04:37הבעיה הגדולה של ביג דאטה: מי יתייג את הכל?
05:31למה תיוג נתונים זה לא פשוט: אשליות אופטיות ופסלים רוקדים
06:32נתונים לא מסומנים, רועשים, רב-ממדיים: מה עושים עכשיו?
07:35פתרון מהפכני: שילוב מסווגים עם מינימום תיוג
09:28לפענח את הקשר: איך היסטוגרמות חושפות מידע הדדי
10:55מאחורי הקלעים: אלגוריתם גרפים דו-צדדיים לניתוח נתונים
12:05הקפיצה הבאה ב-AI: מציאות רציפה והחלטות דיסקרטיות

למרות הרצון להשתמש ביותר נתונים, בעיה מרכזית בביג דאטה היא תיוג הנתונים. קשה לתייג מיליארדי דוגמאות, וגם שיטות כמו למידה בלתי מפוקחת או אשכולות אינן פותרות את הבעיה לחלוטין. האתגר גדל כשהנתונים לא מתויגים, מממדים גבוהים, סדרות זמן ורועשים.

ניתן למפות את כל הנקודות באמצעות שני המסווגים, F1 ו-F2. לאחר מכן, מתייגים מספר קטן של דוגמאות (למשל, 3 מקרים) ומבקשים מכל מסווג למצוא את הדוגמאות הקרובות ביותר לתוויות אלו. לבסוף, בונים היסטוגרמה דו-ממדית כדי לראות אם המסווגים מספקים מידע הדדי, גם אם רוב הנתונים אינם מתויגים.

השיטה מאפשרת להפיק תועלת מנתונים לא מתויגים, ומציעה דרך חדשה לשפר את יכולותינו בתחום. היא מנצלת את העובדה שהעולם רציף אך ההחלטות דיסקרטיות, ומאפשרת להבין טוב יותר את המידע ההדדי בין מסווגים. זה פותח דלתות לתחומים רבים ויכול להוביל לקפיצות משמעותיות בבינה המלאכותית.

גבול ההחלטה בין "טוב" ל"רע" אינו מיושר, בדומה לקבוצת מנדלברוט. פונקציות פרקטליות מראות שלא משנה כמה מתקרבים, תמיד יימצאו נקודות שמתנהגות באופן דומה, וביניהן עולם שלם שמתנהג אחרת. מורכבות זו משקפת את האופי הפרקטלי של המציאות.

ויכוח מוקדם במדעי המחשב עסק בשאלה האם העולם דיסקרטי או רציף. אלן טיורינג טען שהעולם דיסקרטי כי קל יותר לייצג אותו (0 הוא 0, 1 הוא 1). מתנגדיו טענו שהמציאות רציפה ואנלוגית, אך קשה יותר לייצג אותה ולבנות עליה מכשירים.

מייקל רבין הראה שמכונת טיורינג הסתברותית יכולה לבצע משימות יומיומיות, כמו מציאת מספרים ראשוניים, שהם ליבת התקשורת הדיגיטלית. הסתברות חשובה מאוד והיא דרך להפוך את העולם לפחות דיסקרטי ויותר רציף, אך גם הסתברותי.

ראיית מחשב הוא תחום שהחל לפני שנים רבות כפרויקט להוציא מחשבים לעולם. כיום, הוא נבלע במידה רבה בתחום למידת המכונה. ההרצאה סוקרת את התפתחות התחום ואת כיווניו העתידיים, תוך התמקדות באתגרי ייצוג העולם.