Back to IMVC 2026

פענוח סרטונים: ה-AI של NVIDIA שישנה את הכל!

גלו כיצד NVIDIA משנה את עולם ניתוח הווידאו עם פתרונות חיפוש וסיכום מתקדמים. למדו על ארכיטקטורות AI פורצות דרך, זיהוי אנומליות בזמן אמת ויצירת דוחות אוטומטיים. שחררו את הפוטנציאל המלא של נתוני הווידאו שלכם.

Eyal EnavEyal EnavVision AI Alliances Manager, Nvidia

פרקים

00:05מעבר ל-GPU: מהפכת ה-AI של NVIDIA

So I'll be speaking about video search and summarization, video Analytics, Blueprint from Nvidia and VLMs from Nvidia, including Cosmos, which is a foundation model. Not everybody knows that Nvidia is doing much more than GPUs. We're doing also the acceleration for many years now and. Just a second. Okay, yeah, acceleration of all the basic, you know, FFTs, computer vision for many years now, but also doing research, research and creating our own models, own data sets. And in the past months we actually went and distributed that, opened that to the crowd with Cosmos. So I'm coming from a team called Metropolis, Nvidia. Metropolis is focused about smart cities, industrial domain logistics, where the money really, you know, there's a roi, clear roi. Of course, this is relevant to all domains. All the technology of Vision AI is relevant to all domains, including healthcare. And we're also there, we're also very active there. So video search and summarization, who knows about it? Just to let us know if somebody. Okay, great, great, great. So it really brings together all the technology from Nvidia, including computer vision, Classical computer vision Deepstream, which is the pipeline for video analytics, vision language models and agentic AI. This is the latest update. So let's assume you want to process a very long video, okay. And you want to find anomalies, you want to ask questions, you want to get a summarization. That's what video search and summarization does. We have a lot of examples of companies already using that for sequence of operation optimization, of course, Defense security, industrial domain inspection, aoi. Let's see an example to understand better about how it's working. So in this video you can see the nice example of sports analytics. You know, for sports analytics it requires high frame rate, being able to track objects with the obscura. Okay. So with the Nvidia Deepstream and detection and segmentation together with VLMs, with embedding models, we're able to do that and get really all this metadata together into the same database and ask the questions. Okay, ask the questions. And if you want to get alerts in real time, you're able to do that. So you can see it's relevant to all domains. So the basic architecture, it's really classical computer vision ingestion. Very, very high throughput ingestion. Doing decode with multiple GPUs. Okay, this is the first part that is, you know, don't write it on your, on your own with Python or something, or ffmpeg. Just use the ingestion pipelines from Nvidia. This is super fast. Okay, after that we have the Deepstream. Who knows Deepstream. Okay, okay, we need to talk after that. So Deepstream is really the engine for video analytics. Okay. If you want to do video analytics on the edge, it's Deepstream. It's zero copy batching. Everything is accelerated. Everything including the resizing and color conversions and everything. This is the Deepstream and it's part of the ingestion pipeline for VSS. Okay. After that we have VLMs and embedding models. Embedding models like C radio, Dino V3 and others. We as Nvidia, we accelerate all of them of course to get them on tensorrt to be accelerated for best performance and eventually also agentic AI, generic AI today, you know it's, it's the base, right? It's, you know, you have metadata, you can use LLMs to talk to this data. So it's very basic. But we already implemented report generation and deep search. Being able to search anything with the prompt, the agent just breaks the prompt into the relevant elements. So we just will see that in a moment. Okay, so I don't know if you can see that. This is like a more of a zoom in into the design but on the bottom left you see the real time CV element. This is the Deepstream plus detectors, segmentation and tracking. By the way, we have excellent trackers from Nvidia, including 3D trackers. This is really an Nvidia development accelerated. So again, don't go and build your tracker. Of course if you're in a special domain you can do it. But we already have excellent Trackers. Real Time VLM. If you are able to run VLM in Real Time V, there are some lightweight VLMs to find the alert. Okay, you write, I don't know, a person falling, a person running that can be detected in real time. But it's very heavy on compute, right? So you don't want to do it for every frame. So we have also frame selection as an option. After that real time embedding. If you want to do search of features, a person with a red shirt and glasses, that's the features. The C Radio is the best model that we have, but you can plug in your own also and also embedding for video Cosmos embed, video embedding for zero shot video detection. Okay, so if you want to be very fast, you just use those elements without the vlm, the embedding models and the real time computer vision. After that we have behavioral analytics. This is a set of functions like classical machine learning for understanding trajectories, really seeing the anomalies in the metadata. And finally VLM as a verification for the alerts that were detected before. Okay. And if you want to do question and answering summarization and everything. Now for doing the summarization, you need to do sliding window very effectively. We already do it. We're going to release another version that is a VLM that is streamlined for continuous moving window processing. And again, on the top set of capabilities, we released skills skills for the video search and summarization. Every agent today uses skills. You can just use that you can change and create your own skill. Of course, this is really a plug and play. I know many of you are researchers. You can bring in your own models, data pipelines into this thing to really speed up your research too. Okay. It's not only for production. So this is an example here on how it works with Nimo Claw. Who knows Nimo Claw? What? Okay, okay. Nimo Claw is good that you know, nimoklow is an enterprise ready version of openclaw. Okay, you know openclaw that I'm sure. So nimoklow just wraps openclow with the virtual machine and applies policies so you don't get OpenCloud to access your files and network freely. Just protects the OpenCloud. Okay, so that works with the VSS skills and able to do really dynamic search in the video from free text. Okay, so I guess this is becoming of the common, common use case companies going and integrating their own legacy pipelines into agentic pipelines, enabling more flexibilities for the customers. Okay, so here in this example, you see. Let me go back a little bit again. In this example the user just asks what skills do you have? How can I use them? And looking for two players running at the same time. Okay, this is the skills. Okay, I think this is like today it's very common to work in this way and you can build excellent again products and development cycles, development environments for your own internal use too. Okay, so you see the results. It used all the embedding models, the C Radio and the vlm. Now it's working with the VLM to verify, verify those three top results. Okay, and this is using Cosmos 2, we have Cosmos 3 based on Quinn. We take the models and make them commercial. Okay. We train them, we make them

01:00NVIDIA Metropolis: AI לערים חכמות ותעשייה
01:26חיפוש וסיכום וידאו: פתרון ה-VSS של NVIDIA
03:10ארכיטקטורת VSS: מאינג'קשן ל-VLM
04:45ראייה ממוחשבת בזמן אמת ומודלי הטמעה
06:21ניתוח התנהגות ואימות התראות ב-VLM
07:07התאמה אישית של VSS: מיומנויות ומודלים משלכם
07:38Nemo Claw: אבטחה ברמת אנטרפרייז ל-AI
08:44הדגמה חיה: חיפוש אג'נטי בווידאו

good to work with, valid to work with in terms of data and licensing and everything. Okay, let's go a little bit faster. So on the left you can see Agentic Search. Agentic Search is Just use the prompt to use all the elements. The detection, captioning, embedding to find an object, video summarization and report, and real time verified contextual alerts. The search is using all the elements that the VSS is capable of. It's using an LLM to understand the prompt and break it down automatically. Okay. It's a very simple approach, but very efficient. Okay, so here you can see also the pipelines of embedding, separate pipelines for embedding in computer vision with the Cosmos embed and cradle. And after that, using the retrieval of both, you can apply your own metrics to combine both of them. Okay. If you want to combine the computer vision, give more weight to the computer vision or to the embedding, it's your call. Okay, very simple. You can see another example of the actual elements used in real time. Okay, so you see the players. Just a second, let's see. Yeah, so, okay, just get it faster here. So, yeah, once you ask the question, the elements are being activated on the fly. Okay. So the VLM is being called to verify the results. The embedding models. See here The C radio CGLIP2 being used on the fly based on the prompt that was, was used. Okay. And you get all the results and you know, apply your own metrics to it. Okay. Video summarization. Video summarization. You could implement it with the vlm, just calling in a VLM time after time, you know, getting the captions and then activate another LLM on that to get the summarization. But that's not effective, right? You want, eventually you want to be super effective to work with the real systems, minimize the compute costs. Whether you're working in the cloud or on the edge, wherever you're using the compute, you want to minimize the consumption. So we implemented it in very, very efficient way, starting with the ingestion and aggregation of the metadata and everything. Okay. So for reports, you can also define your own templates for reports. Very simple. And get it really tailored for your use. Let's talk about performance a little bit. So I know this is the element that is a little bit, I can say, a bottleneck. Right. Everybody wants to use the biggest vlm, the LLM, in real time in production. But really it's about compute and cost. So you can see that the results here for VSS is really excellent. We have video summarization. You can see the video summarization with a very low latency. This is a 10 minute video summarization. You can see in one minute in RTX 6000 Pro. Okay. You know, this card, this is like a game changer. It has 96 gigabytes of memory and is very cost effective. And a lot of video decoders and encoders. So this is very efficient. The 2L 40s, also very good setup. And the Jetson, you can see the Jetson, Thor and digixpark. Who knows Digic Spark, by the way? Who owns a Digic Spark? Okay, we need to fix that. So Digic Spark is like, you know, the best compute for a researcher because you have 128 gigabytes of memory. Okay. It's like a small digit just, you know, on your desktop. So not. It's a real thing. Okay. So I recommend it. Okay, let's look at another example. So in this case, we're asking for a report. Let's generate a report. Was there some safety issue in the video? Okay. Really cutting edge. So we're just asking free text, the object that we're interested in, the actions that we're interested in, and we get the report. Okay. Very efficiently. Yeah, yeah. So this is, you know, again, you can do it with ChatGPT, with Gemini, you can do it, but this is an on prem solution running very efficiently. Okay. So really speed up your time to market, that's for sure. Okay. Contextual alerts. So in the real world, we want to find things fast, right? We don't want to just process in offline. So for that we really again optimized the complete workflow, the complete pipeline, end to end from the ingestion to the reasoning. You can see the number of streams here. Okay, so for Jetson Thor, you have 14 streams with latency of 09 seconds. It's all, you know, great, great performance, no question. And it uses all the elements, right? You have today you have yolo real time det C, radio, Cosmos embed. And you can plug in your own models into this. Of course, it's better if you have it in VLLM or tensort. It's better for your performance. Okay, let's take a look at an example of real time alerts. So real time alerts, you define the alerts in a free text. Okay. Again, of course. And the blueprint knows how to apply each the alert breakdown to the various elements. Let's say it's a person crossing the line, person running whatever features, in this case a box falling on the ground. Just process it in real time. And you can ask the system to also give it in specific feature like give it actual confidence, the confidence per alert. Okay. And apply more Rules into it. Okay. Performance Overall for alerts, 51 streams with 101 milliseconds. Okay, so this is super fast, super fast. Also for Edge, with a little bit of an effort, you can also squeeze it into a smaller Jetson. Not only Thor, we have a partner here, crg, showing some demos you can see also take a look there and some examples. And eventually again, Nvidia, we are very focused on physical AI. So this is the part, what I showed you is the part of reasoning, right? Understanding videos. But eventually this is really connected to the complete flow of physical AI and robotics. We relate to infrastructure, buildings and cameras and cameras in cities and warehouse as a robot by itself. So you can think about a building set of robots, using models to understand what's happening in real time and also generating predictions in real time for activating actions. Okay, so this is Cosmos, Cosmos from Nvidia. It's really about aggregating everything, consolidating everything together. The reasoning plus the generation plus the action. Action, models, vision, language, action. So we really opened the internal knowledge from Nvidia, where we already created a lot of data sets and models and flow for tuning models. Okay, so we released everything as part of Cosmos. Okay, you can see 15 terabytes of data, 320,000 trajectories, lots of assets for synthetic data, for Omniverse, for the domain where you create synthetic environments. All of this for training robots, testing robots and physical infrastructure. Again, focus on performance. Okay, so Cosmos is a single unified two tower mixture of transformer with two towers trained together, which has its benefits. It has a unified domain that can be used for better understanding. Also for the robot to understand what's happening in real time, predict what is going to happen next and take the action. Let's say an example of. First of all, it takes the lead in the benchmarks, a lot of the benchmarks. And we released, as I said, the physical AI data factory. With this you can go and create your own data and train the models very efficiently. This is an augmentation also for automotive, for robotics and other domains available right now. That's it. That's it for me. I'll be here if you have any questions. Feel free.

10:05יכולות חיפוש אג'נטי: פירוק פרומפטים אוטומטי
12:08סיכום וידאו יעיל: מזעור עלויות חישוב
13:07ביצועים מרשימים: סיכום וידאו בזמן שיא
14:42התראות קונטקסטואליות בזמן אמת ויצירת דוחות
18:01העתיד של AI: בינה פיזית ורובוטיקה עם Cosmos
18:50Cosmos: איחוד חשיבה, יצירה ופעולה
20:33Physical AI Data Factory: מפעל הנתונים לאימון רובוטים
21:11שאלות ותשובות: חיזוי תנועה ואתגרי זמן אמת

Did you include also changing in morphology and kinetics and also forecasting of the motility of the people and the players can hear you. Can I come again? I'm asking if you included morphological and kinetic analysis and forecasting to your metrics. So in the Omniverse, it's already baked in to the models. Okay. Everything is physically grounded. So the data that we train the model with is all physically grounded from simulation itself. Okay. So it includes trajectories and everything. But can you forecast the next step of the football player? Yeah, yeah, yeah. Forecast, you know, multiple scenarios. Right. It's not a single scenario. It can be multiple scenarios based on the frames that we have so far. But this is the actual, you know, today it's not running in real time, you know, the forecast and the action decision. But this is the future. Right. We need the robot to understand what's happening and also forecast multiple scenarios and take action per scenario and then do some another reasoning all in real time. Okay, so this is Cosmos is doing this, but it will take some several steps to make it really real time. We can talk about this also. Thank you. Thank you.

פתרון חיפוש וסיכום וידאו (VSS) של Nvidia מאחד טכנולוגיות שונות של Nvidia, כולל ראייה ממוחשבת, Deepstream, מודלי שפה חזותיים (VLMs) ובינה מלאכותית סוכנת. הוא מאפשר למשתמשים לעבד סרטונים ארוכים כדי למצוא חריגות, לשאול שאלות ולקבל סיכומים. חברות כבר משתמשות בו לאופטימיזציה של רצף פעולות, אבטחה תעשייתית ובדיקות.

Cosmos הוא מודל יסוד של Nvidia המאגד מודלי הסקה, יצירה ופעולה, ומתמקד בבינה מלאכותית פיזית. זהו שילוב אחיד של שני מגדלי טרנספורמרים שאומנו יחד להבנה, חיזוי ופעולה טובים יותר בזמן אמת עבור רובוטים ותשתיות פיזיות. Nvidia פתחה את הידע הפנימי שלה, כולל 15 טרה-בייט של נתונים ו-320,000 מסלולים, כחלק מ-Cosmos לאימון מודלים.

Deepstream הוא מנוע של Nvidia לניתוח וידאו, במיוחד עבור פריסות קצה. הוא מספק צינור עיבוד מואץ עם העתקה אפסית (zero-copy batching) לקליטת וידאו בתפוקה גבוהה, כולל פענוח, שינוי גודל והמרות צבע. Deepstream הוא חלק ליבה מצינור הקליטה של פתרון חיפוש וסיכום וידאו (VSS) של Nvidia.

פתרון VSS של Nvidia מציע ביצועים מצוינים, כאשר סיכום וידאו באורך 10 דקות לוקח כדקה אחת בכרטיס RTX 6000 Pro. עבור התראות מבוססות הקשר, המערכת יכולה לעבד 14 זרמים ב-Jetson Thor עם השהיה של 0.09 שניות, ו-51 זרמים עם השהיה כוללת של 101 מילישניות. יעילות זו מושגת על ידי אופטימיזציה של כל תהליך העבודה מקצה לקצה, מקליטה ועד הסקה.

פתרון VSS של Nvidia מאפשר למשתמשים להגדיר התראות בזמן אמת באמצעות הנחיות טקסט חופשי, אשר התוכנית מפרקת לאחר מכן לאלמנטים רלוונטיים לזיהוי. המערכת מעבדת התראות אלו בזמן אמת, ומספקת תכונות ספציפיות כמו רמת ביטחון לכל התראה. צינור עיבוד קצה-לקצה ממוטב זה, מקליטה ועד הסקה, מאפשר זיהוי מהיר של אירועים כמו נפילת קופסה על הרצפה.

Nimo Claw הוא גרסה מוכנה לארגונים של OpenClaw. הוא עוטף את OpenClaw במכונה וירטואלית ומיישם מדיניות כדי להגן מפני גישה חופשית לקבצים ולרשת. Nimo Claw עובד עם יכולות ה-VSS, ומאפשר חיפוש וידאו דינמי מהנחיות טקסט חופשי תוך הבטחת אבטחה.

Nvidia משתמשת ב-Cosmos כמודל יסוד לבינה מלאכותית פיזית, המקשר הסקה עם מודלי יצירה ופעולה עבור רובוטיקה. היא מתייחסת לתשתיות כמו בניינים ומצלמות כרובוטים בפני עצמם, ומאפשרת להם להבין אירועים בזמן אמת, לחזות תרחישים עתידיים ולהפעיל פעולות. Cosmos מספק מפעל נתונים לבינה מלאכותית פיזית עם מערכי נתונים נרחבים לאימון ובדיקת רובוטים ותשתיות פיזיות.