המודל שיודע לחזות תאונות לפני שהן קורות!
גלו כיצד Nexar משתמשת ב-AI כדי לנתח מיליוני קטעי וידאו מהכביש. פיתחנו מודל פורץ דרך שמזהה סיכונים ותאונות פוטנציאליות בדיוק חסר תקדים. הבינו איך הטכנולוגיה הזו משנה את עתיד הבטיחות בדרכים ופועלת בזמן אמת.

פרקים
Today I'm presenting my colleagues work at nexar, the project that was presented in Denver and cbpr. Last one. And it's around the Nexar driving models or badass. Badass 2.1, 2.0. Sorry. So for those who don't know who's Nexaar. So Nexstar is a dashcam company. It has been around 10 years and so far we've managed to collect a massive amount of data which is around much more than 10 billion miles so far recorded. And we have actively around 300, more than 350 cameras actively daily waiting for the next event to be recorded. Yeah. Okay. And we have very high coverage specifically in the U.S. cameras that are that has intelligence on the edge ready to start recording whenever it identifies any kind of event. Interesting for us. Let's watch a video of nexar. This is only a glimpse of the amount of data and the interesting data that our cameras have successfully recorded. Things are not easily simulated. It exists only in the real data. Very hard to watch. Yeah. So we're talking about a massive amount of data, mainly video data. And we want to train our word model on this data to extract interesting stuff, including all of the our business to satisfy all of our business use cases. And in order to train our word model, we selected V JEPA as an architecture and a training regime. As a model for self supervised, which is unlike the regular self supervised models. It is using masking and prediction on the embedding space and not on the pixel space, which lets it more focus on the dynamics, on the physics, on the events that happening in the video. Unlike the other models that focus on the appearance on this texture, pixels, sharpness of the video, of the image, rather than the dynamics existing. And what it gives is a more efficient model, more capable, it has higher understanding of the physics, it saves all the weights that were dedicated to appearance. It saves them for interpreting and encoding the dynamics. So we took around 20 million clips from Nexar, including all the hard cases that we seen before. We used self supervised learning to train our or to build our word model. And we created three different sizes of it. The first and the top one is the 300 million. With 300 million parameters and using distillation we created the other two, sorry, 90 and 20 million. And from this one we trained it from scratch. Although Meta released their own weights, we did not use them. We had enough data to train from scratch. On top of this word model. We started using few heads on top of it to start downstream tasks and make use of all These encoded videos. So first and what we will focus today is our Badass. It's a model for detecting hazards, risks and future accidents on parallel. We have other projects first to facilitate and ease and have more accuracy in searching video. So our current project using this word model, this word model, we can search much more easily and get extremely interesting data. Instead of based on colors, it would be based on dynamics, based on interesting stuff. And in parallel in the future we are also planning to use this word model in order to plan a future trajectory based on what we are seeing so far. So back to badass. Our first model was already published last year. It's one badass first it wasn't enumerated, now it is badass 1.0. Out of the we took our NAXR world model that we trained, we added very tiny head on top of it and we trained it on a data that we annotated that predicts whether there is a risk situation, whether something that should action should be taken in order to prevent an accident or an event. And training was done only on this head. And the result was astonishing on every one of the available benchmark. This tiny head show a really state of the art result even in the external benchmark and our benchmark. And if you notice the gap on our benchmark is even much higher, it's because our data has higher diversity, more unique cases, more long tail events that wasn't covered in the other benchmarks. So there we can see that Badass Open, which is an open data set that we published on Hugging Face and Badass Pro, which is our model that was trained on our data, achieved best results. And following this success and following our internal analysis of the first model, we decided to train even on a larger amount of data. So badass 1.0 was trained on 400k clips, annotated internally. And for badass 2.0 we decided to enlarge or extend our data to be to cover 2 million. And it's a huge jump, it's about 5x. And in order to get there we used two main annotation processes. First it was the first path, which is the automatic one, almost fully automatic. Using badass one, we annotated randomly yet another set of huge millions of data which is then before being used as a gt they were verified by a human. And then as a second path, actually we after the analysis that we did on battles 1.0 and when we realized the gaps that exist there, we smartly sampled our data. Thanks to our internal search and filtering capabilities, we selected the cases where Bados 1.0 was a bit weaker than expected and we want it to be much more robust in this data. So we sampled these cases. These are referring to sometimes conditions, I mean lighting conditions, the VRUs, some few different aspects of the video that we'll see soon. And given all this data, we. We got badass 2.0. I had a. Should have seen here a video which is not playing. No, no, no. Okay. Help. Question. Okay, maybe we'll see it later. I don't know. We should see a video of betas2. Nevermind. Okay, so betas 2.0 set a new record and the. Actually we can see it here. Yeah, we'll skip the marketing video now focus on the numbers and the bars. So Badass 2.0 shows yet another record in event or risk detection, outperforming all the other. Again, all the other competitives in the different Benchmark scored as one and the gap got even much bigger than 1.0. And specifically where the classes or the sampled event that we took from our big corpus that were intended were specifically collected to fill the gaps that badass 1.0. We looked at it after the training and we evaluated their improvement and they were dramatically improved. So specifically we're talking about detecting events that are related with animals, pedestrians, cyclists and within intersections. So we really identified these weak spots in our previous model. And by targeting these in specific data, we could eventually find an improvement. A huge improvement in the data in the results. Sorry. Now if we look at the graph on the top, we see the performance of all our three different models. Actually didn't mention that. Okay, so from betas 2.0 we created much smaller networks in order to be executed and run on the edge. So we have from betas 2.0 we created 2.0 flash and 2.0 flashlight. The flashlight specifically has only 22 million parameter and if you compare it in weights with Cosmos, even when we trade Cosmos on our data, Flash has 1% of the network. Even with this amount of ordered of magnitudes, Badass 2 flashlight outperformed the Cosmos even when trade on Badass. So the difference is not only in the data, but also on the underlying network and training regime. One thing that we also extract from badass 2.0 is the explainability. So in every video that we stream in the network, not only we get the final score for each small batch of videos, for each small frames of the video, we can also extract from the activations which area of the pixels actually account for the rise or the peak that we saw. In this score. So what we do, we take the long video, we cut it into small clips, each one inferred through the network and we track its signals. And once we see a peak going above some threshold that we know, we see that there's some events going on, we go back to the underlying to the embedding and we find where is the maximum activation happening? We can find it in a square. We can actually find the bounding box that generated this peak. Due to time constraints, I have to go a few moments. That's how we created the distillation of the different models that we have. We were able to launch or to actually deploy badass 2.0 flashlight on the Tor Jetson of Nvidia, running in under 3 milliseconds, which is a huge, much more faster than our first one and can run on real time. And the interesting stuff, the interesting outcome of this project is that we created a model that not only understand what's going on on the road, but even if we take the same model and put it in some really new environment, a warehouse, a sky under the water, it always understands the physics and it also reacts expectedly. We put our model on different environments again, we put it in a warehouse, in the sky, dragons going to collide one into the other. It behaves the same always. It predicts whenever a collision is about to happen. The model did not see at all other data other than the ones we have on the street. And that was the amazing stuff that let us call our model a word model. Thank you.
So in the comparison graphs that you have with the Badass on the top left, you also had the Gemini Badass and you had another model, Badass. What are the difference between your Badass and those Gemini Badass and the second one?
So we created those two variants. We had Cosmos Badass and Gemini Badass by fine tuning and training the different architectures than the one that we adopted with the JEPA in order to compare between the different architectures on same data. And they all lead to the same result, which is V JEPA architecture has outperformed all of them. So we trained on same data, other architectures. Okay. Like for Gemini, it's not like API calls from the LLM or it's like an architecture, we find it using their own, I guess, API. Okay, thanks. Thank you. Thank you.
שאלות ותשובות
Nexar היא חברת מצלמות דרך (dashcam) שפועלת כבר עשר שנים. היא אספה למעלה מ-10 מיליארד מיילים של נתונים ויש לה למעלה מ-350,000 מצלמות פעילות מדי יום, בעיקר בארה"ב. המצלמות מצוידות בבינה מלאכותית בקצה, המאפשרת להן להתחיל להקליט אירועים מעניינים באופן אוטומטי.
"מודל המילים" של Nexar מאומן על כמות עצומה של נתוני וידאו, כולל 20 מיליון קליפים ממצלמות Nexar, כדי לחלץ מידע רלוונטי לצרכים העסקיים שלהם. הוא משתמש בארכיטקטורת V-JEPA ובלימוד בפיקוח עצמי, ואומן מאפס בשלושה גדלים שונים של פרמטרים: 300 מיליון, 90 מיליון ו-20 מיליון.
V-JEPA נבחרה מכיוון שהיא מודל בלמידה בפיקוח עצמי המשתמש במסכות וחיזוי במרחב ההטמעה, ולא במרחב הפיקסלים. גישה זו מאפשרת למודל להתמקד יותר בדינמיקה, בפיזיקה ובאירועים המתרחשים בווידאו, במקום במראה החזותי. זה הופך אותו ליעיל יותר, עם הבנה עמוקה יותר של הפיזיקה.
Badass 2.0 הוא מודל לזיהוי סכנות, סיכונים ותאונות עתידיות, המהווה שיפור משמעותי לגרסה הקודמת. הוא אומן על 2 מיליון קליפים מתויגים, פי חמישה יותר מ-Badass 1.0. המודל הציג שיפור דרמטי בזיהוי אירועים הקשורים לבעלי חיים, הולכי רגל, רוכבי אופניים ובצמתים.
מודל המילים של Nexar מפגין יכולת הכללה יוצאת דופן: הוא מבין את הפיזיקה ומתנהג כמצופה גם בסביבות חדשות לחלוטין, כמו מחסנים, שמיים או מתחת למים. הוא מנבא התנגשויות גם כשלא אומן כלל על נתונים מסביבות אלו, אלא רק על נתוני רחוב. יכולת זו להבין פיזיקה בסביבות שונות היא הסיבה לכך שהוא מכונה "מודל מילים".
Badass 2.0 Flashlight היא גרסה קטנה יותר של המודל, עם 22 מיליון פרמטרים בלבד, המיועדת להפעלה במכשירי קצה. היא נפרסה על Nvidia Jetson ומסוגלת לרוץ בפחות מ-3 מילישניות, מה שמאפשר פעולה בזמן אמת. המודל הזה אף עלה בביצועיו על Cosmos, גם כאשר Cosmos אומן על נתוני Nexar, וזאת למרות שהוא מהווה רק 1% מגודל הרשת של Cosmos.
Badass 2.0 מספק יכולת הסבר בכך שהוא לא רק מציג ציון סיכון סופי לכל וידאו, אלא גם מזהה אילו אזורי פיקסלים תרמו לעלייה או לשיא בציון הסיכון. באמצעות ניתוח האקטיבציות, המודל יכול למצוא תיבת תוחמת המצביעה על האזור הספציפי בו התרחש האירוע שגרם לעליית הסיכון.
עוד מפגשים
הבינה המלאכותית שתציל חיים: מהפכה באבחון רפואי
Idan Bassouk
הסוד ליישור מושלם: למידה עמוקה משנה את פני הראייה הממוחשבת!
Oren Freifeld
לפענח את המציאות: סודות ראיית המחשב והבינה המלאכותית
Yonatan Wexler
AI בפתולוגיה: האם הרדיולוגים באמת מקדימים אותנו ב-20 שנה?
Iris Barshack
הסוד לפתיחת עולם ה-AI: הפסאודו-הופכי הלא-ליניארי
Yamit Ehrlich
האם ישראל תהפוך למעצמת AI עולמית? התפקיד שלך!
Noa Lubin
האם שינוי סדר פשוט יכול לשפר את הקיבוץ שלך ב-77%?
Ofir Lindenbaum
הסוד של מודלי AI: איך לגרום להם לשכוח?
Yehuda Dar