Video AI
584 tools and products — search, filter by language, sort, and open any profile for facts and alternatives.
Nothing matches — clear the search, language filter or sort.
Most adopted 250
opencvC++ · 90,518 ★“Open Source Computer Vision Library”cs-video-courses83,133 ★“List of Computer Science courses with video lectures.”d2l-zhPython · 79,827 ★“《动手学深度学习》:面向中文读者、能运行、可讨论。中英文版被70多个国家的500多所大学用于教学。”AI-For-BeginnersJupyter Notebook · 65,869 ★“12 Weeks, 24 Lessons, AI for All!”ultralyticsPython · 60,801 ★“Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image cla”yolov5Python · 57,908 ★“Ultralytics YOLOv5 in PyTorch for object detection, instance segmentation, classification, training, and expor”ai-engineering-from-scratchPython · 47,327 ★“Learn it. Build it. Ship it for others.”500-AI-Machine-learning-Deep-learning-Computer-vision-NLP-Projects-with-code36,415 ★“500 AI Machine learning Deep learning Computer vision NLP Projects with code”openposeC++ · 34,375 ★“OpenPose: Real-time multi-person keypoint detection library for body, face, hands, and foot estimation”applied-ml30,069 ★“📚 Papers & tech blogs by companies sharing their work on data science & machine learning in production.”d2l-enPython · 29,410 ★“Interactive deep learning book with multi-framework code, math, and discussions. Adopted at 500 universities f”label-studioTypeScript · 28,099 ★“Label Studio is a multi-type data labeling and annotation tool with standardized output format”Pixelle-VideoPython · 27,101 ★“🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine”vit-pytorchPython · 25,488 ★“Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single”pytorch-CycleGAN-and-pix2pixPython · 25,224 ★“Image-to-Image Translation in PyTorch”CVJupyter Notebook · 23,364 ★“✅(已完结)超级全面的 深度学习 笔记【土堆 Pytorch】【李沐 动手学深度学习】【吴恩达 深度学习】【大飞 大模型Agent】”learnopencvJupyter Notebook · 23,080 ★“Learn OpenCV : C++ and Python Examples”gaussian-splattingPython · 22,981 ★“Original reference implementation of "3D Gaussian Splatting for Real-Time Radiance Field Rendering"”MaaAssistantArknightsC++ · 22,674 ★“《明日方舟》小助手,全日常一键长草!| A one-click tool for the daily tasks of Arknights, supporting all clients.”datasetsPython · 21,843 ★“🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulatio”screenpipeRust · 21,134 ★“YC (S26) | Open Computer History | Record your screen continuously locally and provide context to your agents ”video2xC++ · 20,997 ★“A machine learning-based video super resolution and frame interpolation framework. Est. Hack the Valley II, 20”AirSimC++ · 18,413 ★“Open source simulator for autonomous vehicles built on Unreal Engine / Unity, from Microsoft AI & Research”visionPython · 17,872 ★“Datasets, Transforms and Models specific to Computer Vision”instant-ngpCuda · 17,524 ★“Instant neural graphics primitives: lightning fast NeRF and more”Wan2.2Python · 17,219 ★“Wan: Open and Advanced Large-Scale Video Generative Models”Awesome-pytorch-list16,639 ★“A comprehensive list of pytorch related content on github,such as different models,implementations,helper libr”cvatPython · 16,557 ★“Computer Vision Annotation Tool (CVAT) is a leading platform for building high-quality visual datasets for vis”labelmePython · 16,116 ★“Image annotation with Python. Supports polygon, rectangle, circle, line, point, and AI-assisted annotation.”VirgilioJupyter Notebook · 14,947 ★“Your new Mentor for Data Science E-Learning.”DeepLearningExamplesJupyter Notebook · 14,841 ★“State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accura”Duix-AvatarC · 14,741 ★“🚀 Truly open-source AI avatar(digital human) toolkit for offline video generation and digital human cloning.”dlibC++ · 14,432 ★“A toolkit for making real world machine learning and data analysis applications in C++”facenetPython · 14,347 ★“Face recognition using Tensorflow”carlaC++ · 14,307 ★“Open-source simulator for autonomous driving research.”Toonflow-appTypeScript · 14,211 ★“Toonflow 是开源一站式 AI 短剧创作工具,将小说、剧本快速转化为动画短剧。集成 AI 编剧、智能分镜、角色与视频生成,跨平台桌面端轻量部署,助力创作者低成本批量产出视觉内容。Toonflow is an ope”open_clipPython · 14,078 ★“An open source implementation of CLIP.”waoowaooTypeScript · 13,716 ★“首家工业级全流程 AI 影视生产平台。Industry-first professional AI Agent platform for controllable film & video production. Fro”CogVideoPython · 12,964 ★“text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)”pytorch-grad-camPython · 12,954 ★“Advanced AI Explainability for computer vision. Support for CNNs, Vision Transformers, Classification, Object”deep-learning-drizzleHTML · 12,933 ★“Drench yourself in Deep Learning, Reinforcement Learning, Machine Learning, Computer Vision, and NLP by learni”MeshroomPython · 12,911 ★“Node-based Visual Programming Toolbox”CycleGANLua · 12,871 ★“Software that can generate photos from paintings, turn horses into zebras, perform style transfer, and more.”colmapC++ · 12,521 ★“COLMAP - Structure-from-Motion and Multi-View Stereo”CVPR2024-Paper-Code-Interpretation12,470 ★“cvpr2024/cvpr2023/cvpr2022/cvpr2021/cvpr2020/cvpr2019/cvpr2018/cvpr2017 论文/代码/解读/直播合集,极市团队整理”HunyuanVideoPython · 12,444 ★“HunyuanVideo: A Systematic Framework For Large Video Generation Model”ViMaxPython · 12,039 ★“"ViMax: Agentic Video Generation (Director, Screenwriter, Producer, and Video Generator All-in-One)"”nerfstudioPython · 11,911 ★“A collaboration friendly studio for NeRFs”ludwigPython · 11,747 ★“Low-code framework for building custom LLMs, neural networks, and other AI models”segmentation_models.pytorchPython · 11,701 ★“Semantic segmentation models with 500+ pretrained convolutional and transformer-based backbones.”rerunRust · 11,330 ★“Visualize, query, and stream to train on multimodal robotics data.”korniaPython · 11,318 ★“🐍 Geometric Computer Vision Library for Spatial AI”pclC++ · 11,092 ★“Point Cloud Library (PCL)”fiftyoneTypeScript · 11,020 ★“Refine high-quality datasets and visual AI models”openvinoC++ · 10,686 ★“OpenVINO™ is an open source toolkit for optimizing and deploying AI inference”autogluonPython · 10,611 ★“Fast and Accurate ML in 3 Lines of Code”yolov3Python · 10,595 ★“PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, va”caireGo · 10,465 ★“Content aware image resize library”lamaJupyter Notebook · 10,203 ★“🦙 LaMa Image Inpainting, Resolution-robust Large Mask Inpainting with Fourier Convolutions, WACV 2022”X-AnyLabelingPython · 10,137 ★“X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, i”computervision-recipesJupyter Notebook · 9,877 ★“Best Practices, code samples, and documentation for Computer Vision.”U-2-NetPython · 9,849 ★“The code for our newly accepted paper in Pattern Recognition 2020: "U^2-Net: Going Deeper with Nested U-Struct”notebooksJupyter Notebook · 9,621 ★“A collection of tutorials on state-of-the-art computer vision models and techniques. Explore everything from f”RobustVideoMattingPython · 9,490 ★“Robust Video Matting in PyTorch, TensorFlow, TensorFlow.js, ONNX, CoreML!”deeplakeC++ · 9,226 ★“Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling s”rf-detrPython · 9,020 ★“RF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COC”jetson-inferenceC++ · 8,967 ★“Hello AI World guide to deploying deep-learning inference networks and deep vision primitives with TensorRT an”Deep-Learning-Interview-Book8,905 ★“深度学习面试宝典(含数学、机器学习、深度学习、计算机视觉、自然语言处理和SLAM等方向)”SanaPython · 8,792 ★“SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer”introtodeeplearningJupyter Notebook · 8,752 ★“Lab Materials for MIT 6.S191: Introduction to Deep Learning”CompreFaceJava · 8,261 ★“Leading free and open-source face recognition system”ailabC# · 7,854 ★“Experience, Learn and Code the latest breakthrough innovations with Microsoft AI”awesome-object-detection7,503 ★“Awesome Object Detection based on handong1587 github: https://handong1587.github.io/deep_learning/2015/10/09/o”mmagicJupyter Notebook · 7,457 ★“OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI”maths-cs-ai-compendiumTypeScript · 7,338 ★“Become a cracked AI/ML researcher/engineer with this unconventional textbook covering maths, computing, and ML”Final2xTypeScript · 7,303 ★“a cross-platform image super-resolution tool”BackgroundMattingV2Python · 7,189 ★“Real-Time High-Resolution Background Matting”lanceRust · 6,952 ★“Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access”modelsPython · 6,933 ★“Officially maintained, supported by PaddlePaddle, including CV, NLP, Speech, Rec, TS, big models and so on.”awesome-multimodal-ml6,925 ★“Reading list for research topics in multimodal machine learning”pix2pixHDPython · 6,923 ★“Synthesizing and manipulating 2048x1024 images with conditional GANs”donutPython · 6,914 ★“Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generato”daily-paper-computer-vision6,777 ★“记录每天整理的计算机视觉/深度学习/机器学习相关方向的论文”scikit-imagePython · 6,573 ★“Image processing in Python”coursesPython · 6,481 ★“This repository is a curated collection of links to various courses and resources about Artificial Intelligenc”EasyPRC++ · 6,428 ★“(CGCSTCD'2017) An easy, flexible, and accurate plate recognition project for Chinese licenses in unconstrained”smileJava · 6,412 ★“Statistical Machine Intelligence & Learning Engine”awesome-self-supervised-learning6,411 ★“A curated list of awesome self-supervised methods”pytorch-metric-learningPython · 6,338 ★“The easiest way to use deep metric learning in your application. Modular, flexible, and extensible. Written in”vllm-omniPython · 6,197 ★“A framework for efficient model inference with omni-modality models”AI-Job-Notes6,142 ★“AI算法岗求职攻略(涵盖准备攻略、刷题指南、内推和AI公司清单等资料)”opencvsharpC# · 6,068 ★“OpenCV wrapper for .NET”ai-deadlinesJavaScript · 5,997 ★“:alarm_clock: AI conference deadline countdowns”Chinese-CLIPJupyter Notebook · 5,992 ★“Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.”layout-parserPython · 5,773 ★“A Unified Toolkit for Deep Learning Based Document Image Analysis”imagededupPython · 5,665 ★“😎 Finding duplicate images made easy!”video-shotcraftTypeScript · 5,641 ★“AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 152 shot recipe cards, 209 mo”ECCV2022-RIFEPython · 5,564 ★“ECCV2022 - Real-Time Intermediate Flow Estimation for Video Frame Interpolation”sahiPython · 5,472 ★“Framework agnostic sliced/tiled inference + interactive ui + error analysis plots”koharuRust · 5,332 ★“ML-powered manga translator, written in Rust.”sportsPython · 5,312 ★“computer vision and sports”sparrowPython · 5,199 ★“Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM”datascience5,153 ★“This repository is a compilation of free resources for learning Data Science.”mmaction2Python · 5,141 ★“OpenMMLab's Next Generation Video Understanding Toolbox and Benchmark”VideoCrafterPython · 5,072 ★“VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models”Awesome-Transformer-Attention5,048 ★“An ultimately comprehensive paper list of Vision Transformer/Attention, including papers, codes, and related w”BallonsTranslatorPython · 5,046 ★“深度学习辅助漫画翻译工具, 支持一键机翻和简单的图像/文本编辑 | Yet another computer-aided comic/manga translation tool powered by deeplearn”machine_learning_completeJupyter Notebook · 5,036 ★“A comprehensive machine learning repository containing 30+ notebooks on different concepts, algorithms and tec”ml-roadPython · 4,913 ★“Machine Learning and Agentic AI Resources, Practice and Research”deep-person-reidPython · 4,899 ★“Torchreid: Deep learning person re-identification in PyTorch.”pigoGo · 4,728 ★“Fast face detection, pupil/eyes localization and facial landmark points detection library in pure Go.”MaaFrameworkC++ · 4,686 ★“基于图像识别的自动化黑盒测试框架 | An automation black-box testing framework based on image recognition”echomimic_v2Python · 4,642 ★“[CVPR 2025] EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation”neuralangeloPython · 4,612 ★“Official implementation of "Neuralangelo: High-Fidelity Neural Surface Reconstruction" (CVPR 2023)”pysotPython · 4,600 ★“SenseTime Research platform for single object tracking, implementing algorithms like SiamRPN and SiamMask.”PyTorch-Tutorial-2ndJupyter Notebook · 4,582 ★“《Pytorch实用教程》(第二版)无论是零基础入门,还是CV、NLP、LLM项目应用,或是进阶工程化部署落地,在这里都有。相信在本书的帮助下,读者将能够轻松掌握 PyTorch 的使用,成为一名优秀的深度学习工程师。”BEVFormerPython · 4,574 ★“[ECCV 2022] This is the official implementation of BEVFormer, a camera-only framework for autonomous driving p”PINTO_model_zooPython · 4,573 ★“A repository for storing models that have been inter-converted between various frameworks. Supported framework”webotsC++ · 4,566 ★“Webots Robot Simulator”ceres-solverC++ · 4,541 ★“A large scale non-linear optimization library”HunyuanVideo-1.5Python · 4,529 ★“HunyuanVideo-1.5: A leading lightweight video generation model”monodepth2Jupyter Notebook · 4,499 ★“[ICCV 2019] Monocular depth estimation from a single image”AIGC-Interview-Book4,408 ★“【三年面试五年模拟】AIGC/LLM/AI Agent算法工程师面试资源平台。涵盖AIGC、LLM大模型、AI Agent、具身智能、传统深度学习、计算机视觉、自然语言处理、自动驾驶、机器学习、强化学习、大数据挖掘、世界”awesome-data-labeling4,401 ★“A curated list of awesome data labeling tools”lingbot-worldPython · 4,376 ★“Advancing Open-source World Models”lmms-evalPython · 4,371 ★“One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks”d2l-pytorchJupyter Notebook · 4,360 ★“This project reproduces the book Dive Into Deep Learning (https://d2l.ai/), adapting the code from MXNet into ”VLMEvalKitPython · 4,349 ★“Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks”champPython · 4,260 ★“[ECCV 2024] Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance”cs230-code-examplesPython · 4,229 ★“Code examples in pyTorch and Tensorflow for CS230”torchgeoPython · 4,153 ★“TorchGeo: datasets, samplers, transforms, and pre-trained models for geospatial data”open_flamingoPython · 4,119 ★“An open-source framework for training large multimodal models.”Generative-Media-SkillsShell · 4,099 ★“Multi-modal Generative Media Skills for AI Agents (Claude Code, Cursor, Gemini CLI). High-quality image, video”FastVideoPython · 3,990 ★“A unified inference and post-training framework for accelerated video generation.”fast-reidPython · 3,982 ★“SOTA Re-identification Methods and Toolbox”4DGaussiansJupyter Notebook · 3,888 ★“[CVPR 2024] 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering”semantic-routerPython · 3,826 ★“Superfast AI decision making and intelligent processing of multi-modal data.”scenicPython · 3,821 ★“Scenic: A Jax Library for Computer Vision Research and Beyond”dramaclawTypeScript · 3,819 ★“A general-purpose AIGC video engine: script to finished film in one pipeline — dramas, ads, product videos, ot”lightlyPython · 3,795 ★“A python library for self-supervised learning on images.”habitat-simC++ · 3,794 ★“A flexible, high-performance 3D simulator for Embodied AI research.”printfilmJava · 3,771 ★“短剧平台 AI Short Film Motion Comic Generation Platform Industrial AI Motion Comic & Video Workbench”MAGI-1Python · 3,769 ★“MAGI-1: Autoregressive Video Generation at Scale”awesome-industrial-anomaly-detection3,742 ★“Paper list and datasets for industrial image anomaly/defect detection (updating). 工业异常/瑕疵检测论文及数据集检索库(持续更新)。”AI-Engineer-HeadquartersJupyter Notebook · 3,673 ★“A collection of scientific methods, processes, algorithms, and systems to build stories & models.”SageAttentionCuda · 3,660 ★“[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAtte”TurboDiffusionPython · 3,616 ★“TurboDiffusion: 100–200× Acceleration for Video Diffusion Models”face.evoLVePython · 3,588 ★“🔥🔥High-Performance Face Recognition Library on PaddlePaddle & PyTorch🔥🔥”make-senseTypeScript · 3,561 ★“Free to use online tool for labelling photos. https://makesense.ai”LichtFeld-StudioC++ · 3,555 ★“Train, inspect, edit, automate, and export 3D Gaussian Splatting scenes from a single native application.”SiamMaskPython · 3,546 ★“[CVPR19/TPAMI23] SiamMask: A Framework for Fast Online Object Tracking and Segmentation”PersonaLivePython · 3,543 ★“[CVPR 2026] PersonaLive! : Expressive Portrait Image Animation for Live Streaming”ml-courseJupyter Notebook · 3,526 ★“Open Machine Learning course”pytrackingPython · 3,515 ★“Visual tracking library based on PyTorch.”AliceVisionC++ · 3,481 ★“3D Computer Vision Framework”colorizationPython · 3,461 ★“Automatic colorization using deep neural networks. "Colorful Image Colorization." In ECCV, 2016.”anylabelingPython · 3,456 ★“Effortless AI-assisted data labeling with AI support from YOLO, Segment Anything (SAM+SAM2/2.1+SAM3), MobileSA”awesome-hand-pose-estimationPython · 3,387 ★“Awesome work on hand pose estimation/tracking”catalystPython · 3,381 ★“Accelerated deep learning R&D”FCOSPython · 3,346 ★“FCOS: Fully Convolutional One-Stage Object Detection (ICCV'19)”Pyramid-FlowPython · 3,208 ★“[ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling”InternGPTPython · 3,205 ★“InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports”openvino_notebooksJupyter Notebook · 3,205 ★“📚 Jupyter notebook tutorials for OpenVINO™”best_AI_papers_20223,183 ★“A curated list of the latest breakthroughs in AI (in 2022) by release date with a clear video explanation, lin”ComputeLibraryC++ · 3,182 ★“The Compute Library is a set of computer vision and machine learning functions optimised for both Arm CPUs and”MASt3R-SLAMPython · 3,165 ★“[CVPR 2025] MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors”3DDFA_V2Python · 3,148 ★“The official PyTorch implementation of Towards Fast, Accurate and Stable 3D Dense Face Alignment, ECCV 2020.”WebPlotDigitizerJavaScript · 3,144 ★“Computer vision assisted tool to extract numerical data from plot images.”torchscalePython · 3,137 ★“Foundation Architecture for (M)LLMs”mmdeployPython · 3,136 ★“OpenMMLab Model Deployment Framework”VLM_survey3,128 ★“Collection of AWESOME vision-language models for vision tasks”habitat-labPython · 3,104 ★“A modular high-level library to train embodied AI agents across a variety of tasks and environments.”awesome-generative-ai-appsJavaScript · 3,017 ★“50+ open-source generative AI apps you can clone, deploy, and monetize — image generators, video tools, virtua”DynamiCrafterPython · 3,008 ★“[ECCV 2024, Oral] DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors”pythoncode-tutorialsJupyter Notebook · 2,999 ★“The Python Code Tutorials”mAPPython · 2,966 ★“mean Average Precision - This code evaluates the performance of your neural net for object recognition.”MinkowskiEnginePython · 2,955 ★“Minkowski Engine is an auto-diff neural network library for high-dimensional sparse tensors”self-driving-carC++ · 2,923 ★“Udacity Self-Driving Car Engineer Nanodegree projects.”VideoPipeC++ · 2,911 ★“A cross-platform video structuring (video analysis) framework based on CV models & mLLM.”PocketFlowPython · 2,908 ★“An Automatic Model Compression (AutoMC) framework for developing smaller and faster AI applications.”ml-surveys2,902 ★“📋 Survey papers summarizing advances in deep learning, NLP, CV, graphs, reinforcement learning, recommendation”comic-translatePython · 2,898 ★“AI comic and manga translator app/browser extension for automatically translating comics, manga, manhwa, BDs, ”best_AI_papers_20212,894 ★“A curated list of the latest breakthroughs in AI (in 2021) by release date with a clear video explanation, li”computer-vision-in-actionJupyter Notebook · 2,859 ★“A computer vision closed-loop learning platform where code can be run interactively online. 学习闭环《计算机视觉实战演练:算法与”MuseVPython · 2,846 ★“MuseV: Infinite-length and High Fidelity Virtual Human Video Generation with Visual Conditioned Parallel Denoi”DINOPython · 2,830 ★“[ICLR 2023] Official implementation of the paper "DINO: DETR with Improved DeNoising Anchor Boxes for End-to-E”knockknockPython · 2,828 ★“🚪✊Knock Knock: Get notified when your training ends with only two additional lines of code”autodistillPython · 2,758 ★“Images to inference with no labeling (use foundation models to train supervised models).”manga-ocrPython · 2,756 ★“Optical character recognition for Japanese text, with the main focus being Japanese manga”SimpleCVPython · 2,731 ★“The Open Source Framework for Machine Vision”CV-CUDAC++ · 2,717 ★“CV-CUDA™ is an open-source, GPU accelerated library for cloud-scale image processing and computer vision.”LightX2VPython · 2,702 ★“Lightweight Image Video Action Generation Inference Framework”MaaNTEPython · 2,698 ★“MaaNTE. Nevertheless to Everless automatic assistant 异环小助手”norfairPython · 2,677 ★“Lightweight Python library for adding real-time multi-object tracking to any detector.”MimicMotionPython · 2,646 ★“High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance”LongLivePython · 2,557 ★“Long Video Gen Infrastructure”ResearchStudioPython · 2,424 ★“ResearchStudio: Our AI co-author, from research problem to final publication.”youtube-automation-agentJavaScript · 2,372 ★“🎬 Fully automated YouTube channel management with AI agents. Creates, optimizes & publishes videos 24/7. Works”GLM-VPython · 2,366 ★“GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning”ailia-modelsPython · 2,365 ★“The collection of pre-trained, state-of-the-art AI models for ailia SDK”InternVideoPython · 2,362 ★“[ECCV2024] Video Foundation Models & Data for Multimodal Understanding”Matrix-GamePython · 2,307 ★“Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory”temporal-shift-modulePython · 2,222 ★“[ICCV 2019] TSM: Temporal Shift Module for Efficient Video Understanding”HeliosPython · 2,065 ★“Helios: Real Real-Time Long Video Generation Model”Code2VideoPython · 1,997 ★“[ICML 2026] Video generation via code”diffgramPython · 1,909 ★“The AI Datastore for Schemas, BLOBs, and Predictions. Use with your apps or integrate built-in Human Supervisi”awesome-seedance-2-promptsTypeScript · 1,875 ★“🎬 2000+ curated Seedance 2.0 video generation prompts — cinematic, anime, UGC, ads, meme styles. Includes Seed”ReCamMasterPython · 1,852 ★“[ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video”video-search-and-summarizationC++ · 1,817 ★“NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for b”dreamtalkPython · 1,788 ★“Official implementations for paper: DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabi”VideoMAEPython · 1,782 ★“[NeurIPS 2022 Spotlight] VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video P”VBenchPython · 1,742 ★“[CVPR2024 Highlight] VBench - We Evaluate Video Generation”Awesome-CVPR2026-CVPR2025-CVPR2024-CVPR2021-CVPR2020-Low-Level-Vision1,732 ★“A Collection of Papers and Codes for CVPR2026/CVPR2025/CVPR2024/CVPR2021/CVPR2020 Low Level Vision”VideoClawPython · 1,714 ★“🚀 AI 全自动化视频生成员工 | Your First AIGC Coworker. Chat an Idea. Get a Film. 🦞”PaddleVideoPython · 1,703 ★“Awesome video understanding toolkits based on PaddlePaddle. It supports video data annotation tools, lightweig”AutonomousVehicleControlBeginnersGuidePython · 1,643 ★“Python sample codes and documents about Autonomous vehicle control algorithm. This project can be used as a te”video-podcast-makerPython · 1,567 ★“Topic → 4K narrated video for coding agents. v4.0: all TTS via the ttsCN engine component (11 platforms incl. ”MiniMax-MCPPython · 1,563 ★“Official MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech, im”PhantomPython · 1,517 ★“Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment”SALMONN1,510 ★“SALMONN family: A suite of advanced multi-modal LLMs”py-motmetricsPython · 1,487 ★“:bar_chart: Benchmark multiple object trackers (MOT) in Python”yolov4-deepsortPython · 1,428 ★“Object tracking implemented with YOLOv4, DeepSort, and TensorFlow.”video-diffusion-pytorchPython · 1,383 ★“Implementation of Video Diffusion Models, Jonathan Ho's new paper extending DDPMs to Video Generation - in Pyt”TeaCachePython · 1,368 ★“Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model”LocalMiniDramaJavaScript · 1,365 ★“🎬 seedance2接入 开源本地 AI 短剧 & 漫剧生成工具 —— 从故事到成片一站式完成,数据不出本机,短剧工作流管理平台,高灵活度,AI真人剧,AI漫剧本地搞定。 Open-source local AI s”FollowYourPosePython · 1,357 ★“[AAAI 2024] Follow-Your-Pose: This repo is the official implementation of "Follow-Your-Pose : Pose-Guided Text”cosmos-predict2.5Python · 1,351 ★“Cosmos-Predict2.5, the latest version of the Cosmos World Foundation Models (WFMs) family, specialized for sim”MagicTimePython · 1,338 ★“[TPAMI 2025🔥] MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators”ai-fusion-videoJava · 1,326 ★“【融光】 - 基于 Agent 的全流程AI短剧/漫剧/视频创作平台 - Java & agentscope2.0 | Agent-based end-to-end AI creation platform for sh”LancePython · 1,324 ★“A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editi”short-video-makerTypeScript · 1,297 ★“Creates short videos for TikTok, Instagram Reels, and YouTube Shorts using the Model Context Protocol (MCP) an”UNINEXTPython · 1,278 ★“[CVPR'23] Universal Instance Perception as Object Discovery and Retrieval”articulated-animationJupyter Notebook · 1,277 ★“Code for Motion Representations for Articulated Animation paper”StableAvatarPython · 1,256 ★“We present StableAvatar, the first end-to-end video diffusion transformer, which synthesizes infinite-length h”MotusPython · 1,236 ★“Official code of Motus: A Unified Latent Action World Model”FastMOTPython · 1,219 ★“High-performance multiple object tracking based on YOLO, Deep SORT, and KLT 🚀”UniAnimatePython · 1,189 ★“Code for SCIS-2025 Paper "UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animatio”YOLOv8-DeepSORT-Object-TrackingJupyter Notebook · 1,176 ★“YOLOv8 Object Tracking using PyTorch, OpenCV and DeepSORT”gbro-collage-brollPython · 1,168 ★“半调纸拼贴 B-roll 生成 skill:三闸门审批,Gemini Omni Flash 首尾帧组装动画 | Editorial halftone paper-collage B-roll agent skill”MagicDrivePython · 1,166 ★“[ICLR24] Official implementation of the paper “MagicDrive: Street View Generation with Diverse 3D Geometry Con”Yolov5-DeepsortPython · 1,155 ★“最新版本yolov5+deepsort目标检测和追踪,能够显示目标类别,支持5.0版本可训练自己数据集”OC_SORTPython · 1,130 ★“[CVPR2023] The official repo for OC-SORT: Observation-Centric SORT on video Multi-Object Tracking. OC-SORT is ”awesome-grounding1,127 ★“awesome grounding: A curated list of research papers in visual grounding”yolo_rosPython · 1,123 ★“Ultralytics YOLOv8, YOLOv9, YOLOv10, YOLOv11, YOLOv12 for ROS 2”locally-uncensoredTypeScript · 1,100 ★“Plug-and-play local AI studio: uncensored chat, image & video generation, coding agent. Runs abliterated LLMs ”MotionDirectorPython · 1,048 ★“[ECCV 2024 Oral] MotionDirector: Motion Customization of Text-to-Video Diffusion Models.”SCAILPython · 1,040 ★“SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations ”SpargeAttnCuda · 1,030 ★“[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.”3DObjectTrackingC++ · 1,025 ★“Algorithms and Publications on 3D Object Tracking”echomimic_v3Python · 1,024 ★“[AAAI 2026] EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animati”
Established 208
animate-anythingPython · 970 ★“Fine-Grained Open Domain Image Animation with Motion Guidance”Awesome-CV-MasterHub969 ★“:fire: :fire: :fire: A paper list of some recent Computer Vision(CV) works”awesome-3d-4d-world-modelsHTML · 967 ★“[TPAMI 2026] 3D and 4D World Modeling: A Survey”SEINEPython · 967 ★“[ICLR 2024] SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction”videocomposerPython · 957 ★“Official repo for VideoComposer: Compositional Video Synthesis with Motion Controllability”UnicornPython · 951 ★“[ECCV'22 Oral] Towards Grand Unification of Object Tracking”Chat-UniViPython · 941 ★“[CVPR 2024 Highlight🔥] Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and”Causal-ForcingPython · 931 ★“[ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Q”lingbot-videoPython · 927 ★“Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence”FancyVideoPython · 921 ★“Video generation from text&image, 1st-gen”FollowYourClickPython · 908 ★“[AAAI 2025] Follow-Your-Click: This repo is the official implementation of "Follow-Your-Click: Open-domain Reg”VistaPython · 893 ★“[NeurIPS 2024] A Generalizable World Model for Autonomous Driving”ControlVideoPython · 864 ★“[ICLR 2024] Official pytorch implementation of "ControlVideo: Training-free Controllable Text-to-Video Generat”ditto-talkingheadPython · 863 ★“[ACM MM 2025] Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis”UniAnimate-DiTPython · 854 ★“UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer”ConsisIDPython · 853 ★“[CVPR 2025 Highlight🔥] Identity-Preserving Text-to-Video Generation by Frequency Decomposition”SeedVR2846 ★“[ICLR2026] SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training”DiffusionAsShaderPython · 832 ★“[SIGGRAPH 2025] Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control”DiT-ExtrapolationPython · 825 ★“Official implementation for "RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers" (I”lmms-enginePython · 823 ★“A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.”VideoMAEv2Python · 815 ★“[CVPR 2023] VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking”kandinsky-5Python · 808 ★“Kandinsky 5.0: A family of diffusion models for Video & Image generation”Autoregressive-Models-in-Vision-Survey806 ★“[TMLR 2025🔥] A survey for the autoregressive models in vision.”DriveAGIPython · 805 ★“[CVPR 2024 Highlight] GenAD: Generalized Predictive Model for Autonomous Driving”open-webui-toolsPython · 795 ★“Open‑WebUI Tools is a modular toolkit designed to extend and enrich your Open WebUI instance, turning it into ”rcmPython · 783 ★“rCM & Causal-rCM: Leading and Unified Algorithms/Infrastructures for Bidirectional/Autoregressive Video Diffus”awesome-video-generationTeX · 781 ★“A collection of awesome video generation studies.”Matrix-3DPython · 781 ★“Generate large-scale explorable 3D scenes with high-quality panorama videos from a single image or text prompt”InfinityStarPython · 778 ★“[NeurIPS 2025 Oral]Infinity⭐️: Unified Spacetime AutoRegressive Modeling for Visual Generation”HQTrackPython · 753 ★“Tracking Anything in High Quality”mlx-serveZig · 753 ★“Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core”awesome-text-to-video739 ★“A Survey on Text-to-Video Generation/Synthesis.”MagicDrive-V2Python · 723 ★“[ICCV 2025] Official implementation of the paper “MagicDrive-V2: High-Resolution Long Video Generation for Aut”cosmos-transfer2.5Python · 722 ★“Cosmos-Transfer2.5, built on top of Cosmos-Predict2.5, produces high-quality world simulations conditioned on ”nano-world-modelPython · 717 ★“A Minimalist, Batteries-included Repository for Advancing World Model Science.”diffusion-forcing-transformerPython · 707 ★“[ICML 2025] Official PyTorch Implementation of "History-Guided Video Diffusion"”ima2-genTypeScript · 705 ★“Local-first visual generation runtime and studio for people and coding agents, with reproducible image and vid”ChronoEditPython · 704 ★“[ICLR 2026] ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation”SynCamMasterPython · 697 ★“[ICLR'25] SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints”Awesome-Video-World-Models-with-AR-DiffusionTeX · 695 ★“A Curated List of Awesome Video World Models with AR Diffusion: Covering Algorithms, Applications, and Infrast”multi-object-trackerPython · 694 ★“Multi-object trackers in Python”magvit2-pytorchPython · 668 ★“Implementation of MagViT2 Tokenizer in Pytorch”ComfyUI-H3-Motion-ContextPython · 668 ★“Clip chaining for MiniMax H3 in ComfyUI - motion and audio genuinely continue across joins”Flow-FactoryPython · 667 ★“A unified framework for easy reinforcement learning in Flow-Matching models”AlayaRendererPython · 666 ★“Generative World Renderer: an AI-native Renderer for Games and Virtual Worlds.”NOVAPython · 661 ★“[ICLR 2025] Autoregressive Video Generation without Vector Quantization”AnimateLCMPython · 659 ★“[SIGGRAPH ASIA 2024 TCS] AnimateLCM: Computation-Efficient Personalized Style Video Generation without Persona”MiniGPT4-videoPython · 637 ★“Official code for Goldfish model for long video understanding and MiniGPT4-video for short video understanding”comfyui-mcpTypeScript · 617 ★“Local-first, agent-native control plane for ComfyUI — MCP server + sidebar agent that generates images, video ”AetherPython · 608 ★“[ICCV 2025 & ICCV 2025 RIWM Outstanding Paper] Aether: Geometric-Aware Unified World Modeling”SpatialVIDPython · 597 ★“[CVPR 2026] SpatialVID: A Large-Scale Video Dataset with Spatial Annotations”Awesome-RL-for-Video-Generation588 ★“A curated list of papers on reinforcement learning for video generation”ComfyUI_VLM_nodesPython · 585 ★“ComfyUI nodes for vision-language models: Qwen3-VL, Moondream 3, Florence-2, SmolVLM2, InternVL, Gemma 3, Mini”common_metrics_on_video_qualityPython · 585 ★“You can easily calculate FVD, PSNR, SSIM, LPIPS for evaluating the quality of generated or predicted videos.”VEnhancerPython · 578 ★“Official codes of VEnhancer: Generative Space-Time Enhancement for Video Generation”Vista4DPython · 578 ★“Official code, models, and data for Vista4D: Video Reshooting with 4D Point Clouds (CVPR 2026 Highlight)”actionformer_releasePython · 573 ★“Code release for ActionFormer (ECCV 2022)”siamfc-tfPython · 570 ★“SiamFC tracking in TensorFlow.”genblazePython · 560 ★“Genblaze is an open source Python SDK for orchestrating generative AI media pipelines across video, audio, and”Vibe-WorkflowJavaScript · 559 ★“Free, open-source alternative to Weavy AI, Krea Nodes, Freepik Spaces & FloraFauna AI — node-based AI workflow”VideoTunaPython · 553 ★“Let's finetune video generation models!”ai-shortVideo-pipelinePython · 552 ★“End-to-end AI short-video production pipeline. FastAPI orchestration + Spring Boot gateway with multi-model fa”Rectlabel-supportJupyter Notebook · 552 ★“RectLabel is an offline image annotation tool for object detection and segmentation.”FreeInitPython · 544 ★“[ECCV 2024] FreeInit: Bridging Initialization Gap in Video Diffusion Models”motpyPython · 538 ★“Library for tracking-by-detection multi object tracking implemented in python”forge-filmPython · 534 ★“Multi-model DAG-driven parallel AI film generation — parallel speedup scales with scene independence; Generate”storytellerPython · 534 ★“Multimodal AI Story Teller, built with Stable Diffusion, GPT, and neural text-to-speech”flatkey-cliJavaScript · 527 ★“Flatkey media generation CLI for images, videos, audio, text, credits, and model discovery.”MeViSPython · 526 ★“[ICCV 2023 & TPAMI 2025] MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions”Cosmos-Drive-DreamsJupyter Notebook · 526 ★“Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models”fantasy-portraitPython · 508 ★“FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers”DiffMorpherPython · 507 ★“Official Code for DiffMorpher: Unleashing the Capability of Diffusion Models for Image Morphing (CVPR 2024)”pytorch-vsumm-reinforcePython · 504 ★“Unsupervised video summarization with deep reinforcement learning (AAAI'18)”siam-motPython · 499 ★“SiamMOT: Siamese Multi-Object Tracking”Open-OmniVCusPython · 498 ★“OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions (NeurIPS 2025)”UpscalerShell · 493 ★“A consolidation of various compiled open-source AI image/video upscaling product for a working CLI friendly im”DAM4SAMPython · 492 ★“[CVPR 2025, IJCV 2026] "A Distractor-Aware Memory for Visual Object Tracking with SAM2", "Distractor-Aware Mem”floatPython · 490 ★“[ICCV 2025] Official Pytorch Implementation of FLOAT: Generative Motion Latent Flow Matching for Audio-driven ”OmniWorldPython · 490 ★“[ICLR 2026] OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling”telememPython · 484 ★“TeleMem is a high-performance drop-in replacement for Mem0, featuring semantic deduplication, long-term dialog”FlashVideoPython · 484 ★“[AAAI-2026]FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation”LV-DOTC++ · 480 ★“LV-DOT: LiDAR-Visual Dynamic Obstacle Detection and Tracking (C++/Python/ROS)”CVPR23_LFDMPython · 470 ★“The pytorch implementation of our CVPR 2023 paper "Conditional Image-to-Video Generation with Latent Flow Diff”Fast-Powerful-Whisper-AI-Services-APIPython · 469 ★“⚡ 一款用于自动语音识别 (ASR)、翻译的高性能异步 API。不需要购买Whisper API,使用本地运行的Whisper模型进行推理,并支持多GPU并发,针对分布式部署进行设计。还内置了包括TikTok、抖音等社交”EDTalkPython · 468 ★“[ECCV 2024 Oral] EDTalk - Official PyTorch Implementation”video-recap-skillsPython · 467 ★“Clip any video into a narration recap with claude code skill|用 claude code skill 把任何视频剪辑成中文解说视频,支持剪映导出”OmniShowPython · 463 ★“[ICML 2026] ByteDance's All-in-One Video Generation Model for Human-Object Interaction Video Generation”ano_pred_cvpr2018Python · 462 ★“Official implementation of Paper Future Frame Prediction for Anomaly Detection -- A New Baseline, CVPR 2018”Awesome-Evaluation-of-Visual-Generation461 ★“A list of works on evaluation of visual generation models, including evaluation metrics, models, and systems”ultralytics-YOLO-DeepSort-ByteTrack-PyQt-GUIPython · 459 ★“a GUI application, which uses YOLOs (YOLOv8, YOLO11, YOLOv13) for Object Detection/Tracking, Human Pose Estima”Seedance-2.5-APIPython · 457 ★“Python wrapper for ByteDance's Seedance 2.5 API — Text-to-Video, Image-to-Video, realistic human faces, native”Video-RAG-masterPython · 455 ★“✨✨[NeurIPS 2025] This is the official implementation of our paper "Video-RAG: Visually-aligned Retrieval-Augme”AnimeInterpPython · 455 ★“The code for CVPR21 paper "Deep Animation Video Interpolation in the Wild"”MM-DiffusionPython · 451 ★“[CVPR'23] MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video Generation”RefAlignPython · 449 ★“[ECCV 2026] Official PyTorch implementation of RefAlign: Representation Alignment for Reference-to-Video Gener”Video-As-PromptPython · 448 ★“[ICLR 2026] Official repo for paper "Video-As-Prompt: Unified Semantic Control for Video Generation"”waifuExtensionSwift · 448 ★“The waifu2x & Other image-enlargers on Mac”MOSS-VLPython · 446 ★“MOSS-VL is the core multimodal model series within the OpenMOSS ecosystem, dedicated to visual understanding.”vmemPython · 446 ★“[ICCV 2025 ⭐highlight⭐] Implementation of VMem: Consistent Interactive Video Scene Generation with Surfel-Inde”Awesome-Try-On-Models444 ★“A repository for organizing papers, codes and other resources related to Virtual Try-on Models”sn-gamestatePython · 442 ★“[CVPRW'24] SoccerNet Game State Reconstruction: End-to-End Athlete Tracking and Identification on a Minimap (C”falldetection_openpifpafPython · 431 ★“Fall Detection using OpenPifPaf's Human Pose Estimation model”wind-comicTypeScript · 428 ★“Multi-agent AI pipeline that turns one line of text into a finished short-form drama: script, cinematic storyb”hyperframes-motion-directorJavaScript · 423 ★“Agent Skill for Chinese-first HyperFrames motion-video production from articles, products, websites, and READM”World-R1Python · 417 ★“[ICML 2026] World-R1: Reinforcing 3D Constraints for Text-to-Video Generation”OpenDWMPython · 416 ★“An open source code repository of driving world models, with training, inferencing, evaluation tools, and pret”lidar_obstacle_detectorC++ · 406 ★“3D LiDAR Object Detection & Tracking using Euclidean Clustering, RANSAC, & Hungarian Algorithm”TesserActPython · 405 ★“ICCV 2025 | TesserAct: Learning 4D Embodied World Models”unified_video_actionPython · 405 ★“Official PyTorch Implementation of Unified Video Action Model (RSS 2025)”Awesome-Generation-Acceleration403 ★“📚 Collection of awesome generation acceleration resources.”MaestroPython · 396 ★“An all-in-one, 100% local AI video, image, and music studio. Director mode plans full music videos and short f”NexiorVue · 392 ★“Consumer AI app for chat, image generation, video generation, and music creation powered by Ace Data Cloud API”stylegan-vPython · 392 ★“[CVPR 2022] StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2”ARIS-in-AI-OfferPython · 391 ★“Bilingual (中文+EN) ML / LLM / diffusion / agent interview cheat sheets for AI 秋招 — generated by ARIS /interview”Transformer_Tracking389 ★“This repository is a paper digest of Transformer-related approaches in visual tracking tasks.”TDNPython · 386 ★“[CVPR 2021] TDN: Temporal Difference Networks for Efficient Action Recognition”World-Simulator383 ★“[IEEE TPAMI 2026] Simulating the Real World: Survey & Resources, which contains our survey "Simulating the Rea”EponaPython · 382 ★“Official Code for Epona: Autoregressive Diffusion World Model for Autonomous Driving (ICCV 2025)”JavisDiTPython · 381 ★“[ICLR 2026] Official implementation of JavisDiT and JavisDiT++ series.”UniVTGPython · 379 ★“[ICCV 2023] UniVTG: Towards Unified Video-Language Temporal Grounding”TrafficLab-3DPython · 378 ★“Create a digital-twin style traffic visualization using only mp4 CCTV footage and its Google Maps location.”openOiiPython · 376 ★“故事想法 → 多智能体协作 → 漫剧成片 | 基于 LangGraph 的 AI 漫剧生成平台”mega-data-factoryPython · 375 ★“🏭 Mega Scale Multimodal DataPipeline for SOTA Foundation Models”UCMCTrackPython · 374 ★“[AAAI 2024] UCMCTrack: Multi-Object Tracking with Uniform Camera Motion Compensation. UCMCTrack achieves SOTA”SpecVQGANJupyter Notebook · 372 ★“Source code for "Taming Visually Guided Sound Generation" (Oral at the BMVC 2021)”yolov8-object-trackingPython · 372 ★“YOLOv8 Object Tracking Using PyTorch, OpenCV and Ultralytics”higgsfield-ai-prompt-skillPython · 371 ★“Claude AI skill for cinematic Higgsfield AI prompts — 32 sub-skills covering Seedance 2.5 (omni-reference, vid”Vidu4DPython · 370 ★“[TPAMI 2025, NeurIPS 2024] Video4DGen: Enhancing Video and 4D Generation through Mutual Optimization”lanshu-awesome-ai-video-kitHTML · 363 ★“做企业 AI 视频项目逼出来的工具包 · 411 prompt · 15 模型 · 7 Claude Skill · 14 篇方法论”large-performance-model.github.ioHTML · 362 ★“LPM 1.0: Video-based Character Performance Model”VideoScenePython · 353 ★“[CVPR 2025 Highlight] VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step”physgenPython · 353 ★“PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation (ECCV 2024)”MA-LMMPython · 352 ★“(2024CVPR) MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding”PixelReferJupyter Notebook · 352 ★“The code for PixelRefer & VideoRefer”CAINPython · 351 ★“Source code for AAAI 2020 paper "Channel Attention Is All You Need for Video Frame Interpolation"”HumanVidPython · 349 ★“[NeurIPS D&B Track 2024] Official implementation of HumanVid”IFRNetPython · 346 ★“IFRNet: Intermediate Feature Refine Network for Efficient Frame Interpolation (CVPR 2022)”OpenTADPython · 345 ★“OpenTAD is an open-source temporal action detection (TAD) toolbox based on PyTorch.”HunyuanPortraitPython · 345 ★“[CVPR-2025] The official code of HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation”Seedance-2-APIPython · 341 ★“Python wrapper for ByteDance's Seedance 2.0 , Seedance 2.5 and Seedance 2 Mini API — Text-to-Video, Image-to-V”RealVideoPython · 339 ★“A real-time streaming conversational video system that transforms text interactions into continuous, high-fide”ai-video-generator-claudePython · 337 ★“10 Claude skills that generate studio-quality AI video prompts for Seedance 2.0 on Higgsfield. Viral hooks, Sa”VIAMEPython · 336 ★“Video and Image Analytics for Multiple Environments”btrackPython · 335 ★“Bayesian multi-object tracking”sdkTypeScript · 334 ★“AI video generation SDK — JSX for videos. One API for Kling, Flux, ElevenLabs, Veed. Built on Vercel AI SDK.”SLAPython · 334 ★“SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear Attention”AstraPython · 333 ★“[ICLR 2026] Astra : General Interactive World Model with Autoregressive Denoising"”CenterPointC++ · 332 ★“TensorRT deployment for CenterPoint Lidar Detection Model.”physics-IQ-benchmarkPython · 330 ★“Benchmarking physical understanding in generative video models”MagiCompilerPython · 330 ★“A plug-and-play compiler that delivers free-lunch optimizations for both inference and training.”Awesome-Physics-Cognition-based-Video-Generation326 ★“A comprehensive list of papers investigating physical cognition in video generation, including papers, codes, ”OmniTokenizerPython · 326 ★“[NeurIPS 2024]OmniTokenizer: one model and one weight for image-video joint tokenization.”UltimateLabelingPython · 325 ★“A multi-purpose Video Labeling GUI in Python with integrated SOTA detector and tracker”VolleyVisionPython · 323 ★“Applying Deep Learning Approaches to Volleyball Data”PVDMPython · 322 ★“[CVPR'23] Video Probabilistic Diffusion Models in Projected Latent Space”DirectorsConsolePython · 320 ★“A web application for prompt generation and multiple ComfyUI remote and local connections for image and video ”XVFIPython · 311 ★“[ICCV 2021, Oral 3%] Official repository of XVFI”pharPython · 309 ★“deep learning sex position classifier”YOLOv8_Segmentation_DeepSORT_Object_TrackingJupyter Notebook · 294 ★“YOLOv8 Segmentation with DeepSORT Object Tracking (ID + Trails)”MicroLensPython · 292 ★“A Large Short-video Recommendation Dataset with Raw Text/Audio/Image/Videos (Talk Invited by DeepMind).”FluidFramesPython · 274 ★“FluidFrames | video AI frame-generation app”vidsumPython · 271 ★“Generate summary of any video :tv: anywhere and anytime”AMTPython · 271 ★“Official code for "AMT: All-Pairs Multi-Field Transforms for Efficient Frame Interpolation" (CVPR2023)”ByteTrack-cppC++ · 269 ★“C++ implementation of ByteTrack that does not include an object detection algorithm.”human-action-classificationPython · 267 ★“Human action classification system with pose-based (MediaPipe) and video-based (3D CNN) models. Features 100+ ”SimilariRust · 266 ★“A framework for building high-performance real-time multiple object trackers”padel_analyticsPython · 265 ★“AI-powered padel analytics”MultiversePython · 260 ★“Dataset, code and model for the CVPR'20 paper "The Garden of Forking Paths: Towards Multi-Future Trajectory Pr”TrackNet-Badminton-Tracking-tensorflow2Python · 256 ★“TrackNet for badminton tracking using tensorflow2”Cap4VideoPython · 254 ★“【CVPR'2023 Highlight & TPAMI】Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?”QD-DETRPython · 251 ★“Official pytorch repository for "QD-DETR : Query-Dependent Video Representation for Moment Retrieval and Highl”Person-Detection-and-TrackingPython · 250 ★“A tensorflow implementation with SSD model for person detection and Kalman Filtering combined for tracking”TensorRT-YOLOv8-ByteTrackC++ · 249 ★“An object tracking project with YOLOv8 and ByteTrack, speed up by C++ and TensorRT.”Awesome-ICCV2025-ICCV2021-Low-Level-Vision249 ★“A Collection of Papers and Codes for ICCV2025/ICCV2021 Low Level Vision”Awesome-Computer-Vision248 ★“Awesome Resources for Advanced Computer Vision Topics”Awesome-Low-Level-Vision-Research-Groups247 ★“A Collection of Low Level Vision Research Groups”TAdaConvPython · 246 ★“[ICLR 2022] TAda! Temporally-Adaptive Convolutions for Video Understanding. This codebase provides solutions f”vidatVue · 242 ★“Video Annotation Tool”TeViTPython · 241 ★“Temporally Efficient Vision Transformer for Video Instance Segmentation, CVPR 2022, Oral”SceneSegPython · 239 ★“Codebase for CVPR2020 A Local-to-Global Approach to Multi-modal Movie Scene Segmentation”LinferC++ · 235 ★“基于TensorRT的C++高性能推理库,Yolov10, YoloPv2,Yolov5/7/X/8,RT-DETR,单目标跟踪OSTrack、LightTrack。”notebooksJupyter Notebook · 233 ★“Ultralytics YOLO tutorials for Colab, Kaggle, and SageMaker covering training, inference, export, and vision t”hawkPython · 224 ★“🔥[NeurIPS 2024] Official Implementation of Hawk: Learning to Understand Open-World Video Anomalies”Computer-Vision-ProjectsJupyter Notebook · 222 ★“All Computer Vision Projects - Beginner to Advanced”Awesome-Open-Vocabulary-Detection-and-Segmentation220 ★“Awesome OVD-OVS - A Survey on Open-Vocabulary Detection and Segmentation: Past, Present, and Future”DeepStreamC++ · 214 ★“NVIDIA DeepStream Monorepo: DeepStream SDK and reference apps for building GPU‑accelerated, real-time video an”VLM-AutoYOLOPython · 214 ★“AI Auto Annotation & YOLO Training Pipeline, End-to-end object detection auto-labeling and YOLO training platf”VidChaptersJupyter Notebook · 213 ★“[NeurIPS 2023 D&B] VidChapters-7M: Video Chapters at Scale”video-SALMONN-2Python · 209 ★“video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual”Text4VisPython · 199 ★“【AAAI'2023 & IJCV】Transferring Vision-Language Models for Visual Recognition: A Classifier Perspective”TubeDETRPython · 194 ★“[CVPR 2022 Oral] TubeDETR: Spatio-Temporal Video Grounding with Transformers”simtrackPython · 192 ★“Exploring Simple 3D Multi-Object Tracking for Autonomous Driving (ICCV 2021)”AlphaRefinePython · 191 ★“Official implementation for the CVPR2021 paper Alpha-Refine”NExT-QAPython · 190 ★“NExT-QA: Next Phase of Question-Answering to Explaining Temporal Actions (CVPR'21)”MMPD_rPPG_datasetPython · 183 ★“MMPD: Multi-Domain Mobile Video Physiology Dataset(EMBC2023 Oral)”Shot2StoryPython · 180 ★“A new multi-shot video understanding benchmark Shot2Story with comprehensive video summaries and detailed shot”PeekingDuckPython · 179 ★“A modular framework built to simplify Computer Vision inference workloads.”football-analysisJupyter Notebook · 176 ★“A comprehensive tool for processing and analyzing video footage, producing detailed insights into gameplay and”awesome-video-self-supervised-learningHTML · 173 ★“A curated list of awesome self-supervised learning methods in videos”Awesome-ECCV2022-Low-Level-Vision170 ★“A Collection of Papers and Codes in ECCV2022 about low level vision”SparseTrackPython · 167 ★“Official PyTorch implementation of SparseTrack”object-tracking-line-crossing-area-intrusionPython · 167 ★“Deep learning based object tracking with line crossing and area intrusion detection”FrozenBiLMPython · 159 ★“[NeurIPS 2022] Zero-Shot Video Question Answering via Frozen Bidirectional Language Models”BIKEPython · 156 ★“【CVPR'2023】Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Langu”ST-LLMPython · 155 ★“[ECCV 2024🔥] Official implementation of the paper "ST-LLM: Large Language Models Are Effective Temporal Learne”CGDETRPython · 154 ★“Official pytorch repository for CG-DETR "Correlation-guided Query-Dependency Calibration in Video Representati”video2tfrecordPython · 153 ★“Easily convert RGB video data (e.g. .avi) to the TensorFlow tfrecords file format for training e.g. a NN in Te”paper-listPython · 152 ★“autoupdate paper list”
More tools 126
auroraPython · 147 ★“[ICLR 2025] AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark”HolmesVAUPython · 144 ★“[CVPR 2025 Highlight] Official implementation of "Holmes-VAU: Towards Long-term Video Anomaly Understanding at”long-short-term-transformerPython · 140 ★“[NeurIPS 2021 Spotlight] Official implementation of Long Short-Term Transformer for Online Action Detection”AVIONPython · 138 ★“[arXiv:2309.16669] Code release for "Training a Large Video Model on a Single Machine in a Day"”AchelousPython · 137 ★“The official repository of Achelous and Achelous++”mvdPython · 135 ★“[CVPR2023] Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representat”DEARPython · 133 ★“[ICCV 2021 Oral] Deep Evidential Action Recognition”SynchformerPython · 132 ★“Source code for "Synchformer: Efficient Synchronization from Sparse Cues" (ICASSP 2024)”QTSplusPython · 131 ★“Query-aware Token Selector (QTSplus), a lightweight yet powerful visual token selection module that serves as ”DecAlignJavaScript · 131 ★“[ICLR 2026] DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning”rpi-urban-mobility-trackerJupyter Notebook · 130 ★“The easiest way to count pedestrians, cyclists, and vehicles on edge computing devices or live video feeds.”just-askJupyter Notebook · 128 ★“[ICCV 2021 Oral + TPAMI] Just Ask: Learning to Answer Questions from Millions of Narrated Videos”AiATrackPython · 128 ★“[ECCV'22] The official PyTorch implementation of our ECCV 2022 paper: "AiATrack: Attention in Attention for Tr”yolo_series_deepsort_pytorchPython · 126 ★“Deepsort with yolo series. This project support the existing yolo detection model algorithm (YOLOV8, YOLOV7, ”tracktorJupyter Notebook · 125 ★“Python and OpenCV based object tracking software”top-100-computer-vision-projects-idea-for-2024125 ★“Welcome to the "Top 100 Computer Vision Projects Idea for 2024" repository! This repository contains a curated”TAO-AmodalPython · 123 ★“Official Code for Tracking Any Object Amodally”YOLOV8-DeepSORT-Tracking-Vehicle-CountingPython · 123 ★“Vehicle Counting Using Yolov8 and DeepSORT”yolov5-object-trackingPython · 123 ★“YOLOv5 Object Tracking + Detection + Object Blurring + Streamlit Dashboard Using OpenCV, PyTorch and Streamlit”LLaVA-ScissorPython · 122 ★“The official code for the paper: LLaVA-Scissor: Token Compression with Semantic Connected Components for Video”AWTPython · 121 ★“[NeurIPS 2024] AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation”FCSNPython · 117 ★“A PyTorch reimplementation of FCSN in paper "Video Summarization Using Fully Convolutional Sequence Networks"”STALEPython · 116 ★“[ECCV 2022] Official Pytorch Implementation of the paper : " Zero-Shot Temporal Action Detection via Vision-La”volleyball_analyticsPython · 113 ★“This project is designed to display how we can utilize deep learning methods for Sports Data Analytics.”CVPR2024-FACTPython · 107 ★“Official Repo for CVPR 2024 Paper "FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Fully-Su”Fitness-AQAPython · 106 ★“Fitness Action Quality Assessment or your AI-Fitness Coach [ECCV 2022]”diveTypeScript · 105 ★“Media annotation and analysis tools for web and desktop. Get started at https://viame.kitware.com”OpenPVSGJupyter Notebook · 104 ★“Benchmarking Panoptic Video Scene Graph Generation (PVSG), CVPR'23”WorldMMPython · 104 ★“[CVPR 2026 Highlight] WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning”YOLOv7-DeepSORT-Object-TrackingJupyter Notebook · 104 ★“YOLOv7 Object Tracking using PyTorch, OpenCV and DeepSORT”Extended-Kalman-FilterC++ · 104 ★“Implementation of an EKF in C++”demo2programPython · 103 ★“An official TensorFlow implementation of "Neural Program Synthesis from Diverse Demonstration Videos" (ICML 20”WebAR.rocks.trainJavaScript · 102 ★“Object detection, tracking, and 6DoF pose estimation in the web browser, integrated training environment to tr”robotics-level-4Python · 99 ★“This repo contains projects created using TensorFlow-Lite on Raspberry Pi and Teachable Machine. AI and ML ca”yolov9-bytetrack-tensorrtC++ · 98 ★“Integration of YOLOv9 with ByteTracker”ReflectWorldTypeScript · 95 ★“ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Stream”WebAR.rocks.objectJavaScript · 95 ★“Lightweight WebGL and JavaScript library for real time object detection, tracking and 6DoF pose estimation in ”MMNPython · 91 ★“[AAAI 2022] Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding”visionOS-2-Object-Tracking-DemoSwift · 91 ★“visionOS 2 + Object Tracking + ARKit means: we can create visual highlights of real world objects around us an”VideoHighlighterPython · 89 ★“Open-source local AI video analyzer powered by Ollama. Visual search, automatic highlights, scene/action/objec”context-aware-ragPython · 89 ★“Context-Aware RAG library for Knowledge Graph ingestion and retrieval functions.”Fast-TrackPython · 88 ★“Object tracking pipelines complete with RF-DETR, YOLOv9, YOLO-NAS, YOLOv8, and YOLOv7 detection and BYTETracke”RacketVisionPython · 88 ★“The official repository of the paper "RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Ra”Awesome-Multimodal-Reasoning87 ★“Latest Advances on (RL based) Multimodal Reasoning and Generation in Multimodal LLMs”GRMPython · 87 ★“[CVPR'23] The official PyTorch implementation of our CVPR 2023 paper: "Generalized Relation Modeling for Trans”FrameShiftC# · 87 ★“Offline media processing utility powered by FFmpeg, local AI, and fast right-click workflows.”SDNPython · 86 ★“[NeurIPS 2019] Why Can't I Dance in the Mall? Learning to Mitigate Scene Bias in Action Recognition”advanced-computer-vision-engineer-roadmap-202485 ★“A comprehensive roadmap that outlines the key steps and topics you should cover on your journey to becoming a ”YOLOv9_DeepSORTJupyter Notebook · 83 ★“This repository contains code for object detection and tracking in videos using the YOLOv9 object detection mo”OpenCV-Object-Tracker-Python-SamplePython · 82 ★“Python版OpenCVのTracking APIの比較サンプル”Impossible-VideosPython · 81 ★“ICML 2025 - Impossible Videos”Video-BenchPython · 80 ★“Video Generation Benchmark”YOLOv8-Object-Detection-Tracking-Image-Segmentation-Pose-EstimationPython · 80 ★“YOLOv8 object detection, tracking, image segmentation and pose estimation app using Ultralytics API (for detec”QuoTAPython · 79 ★“✨✨[AAAI 2026] This is the official implementation of our paper "QuoTA: Query-oriented Token Assignment via CoT”MTL-AQAPython · 77 ★“What and How Well You Performed? A Multitask Learning Approach to Action Quality Assessment [CVPR 2019]”FluxMemPython · 76 ★“[CVPR 2026] FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding”FISRPython · 76 ★“[AAAI 2020] Official repository of FISR.”multi-camera-people-trackingPython · 75 ★“Multi-camera people tracking is to monitoring people with multiple cameras and connecting each other”Awesome-Video-MLLMs71 ★“:fire: :fire: :fire: Awesome MLLMs/Benchmarks for Short/Long/Streaming Video Understanding :video_camera:”Mask4FormerPython · 71 ★“Mask4Former: Mask Transformer for 4D Panoptic Segmentation”CSTAPython · 70 ★“The official code of "CSTA: CNN-based Spatiotemporal Attention for Video Summarization"”VideoMAE-Action-DetectionPython · 70 ★“[NeurIPS 2022 Spotlight] VideoMAE for Action Detection”GLUSJupyter Notebook · 70 ★“[CVPR 2025] Official PyTorch Implementation of GLUS: Global-Local Reasoning Unified into A Single Large Langua”Track-trtC++ · 70 ★“基于 TensorRT 的 C++ 高性能单目标跟踪推理,支持算法OSTrack、LightTrack。”FreeVAPython · 69 ★“FreeVA: Offline MLLM as Training-Free Video Assistant”SquotSmalltalk · 69 ★“Squeak Object Tracker - Version control for arbitrary objects, currently with Git storage”Animation-from-BlurPython · 69 ★“[ECCV2022] Animation from Blur: Multi-modal Blur Decomposition with Motion Guidance”SJTU-VideoAnalysisJupyter Notebook · 68 ★“智能视频分析:视频目标检测,视频人群计数”VideoLucyPython · 68 ★“[NeurIPS 2025] Deep Memory Backtracking for Long Video Understanding”OmniAgentPython · 67 ★“OmniAgent (ICML 2026): the first native omni-modal agent for active video perception — a 7B agent that beats Q”fast-volleyball-tracking-inferencePython · 67 ★“Fast Volleyball Tracking Inference: Real-time volleyball ball detection and tracking at 100 FPS on CPU (Intel ”byte-track-eigenC++ · 67 ★“ByteTrack-Eigen is a C++ implementation of the ByteTrack object tracking method, leveraging the Eigen library ”SFSORTJupyter Notebook · 66 ★“SFSORT: Scene Features-based Simple Online Real-Time Tracker”R1-TrackPython · 66 ★“R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning.”DIN-Group-Activity-Recognition-BenchmarkPython · 65 ★“[ICCV 2021] A new codebase containing various methods for Group Activity Recognition. Paper title: Spatio-Temp”Object-tracking-and-counting-using-YOLOV8Jupyter Notebook · 65 ★“This repository contains the code for an object detection, tracking and counting project using the YOLOv8 obje”UAV-Drone-Object-Tracking-using-Kalman-FilterPython · 64 ★“This project proposes the implementation of a Linear Kalman Filter from scratch to track stationary objects an”ai-video-summarizerPython · 64 ★“An AI-powered tool for transcribing, summarizing, and creating smart clips from video and audio content.”teresRust · 64 ★“🎞️ Utility for realistic motion blur through frame intepolation and blending”YOLOv8-DeepSORT-StreamlitPython · 63 ★“YOLOv8 Object Tracking and Counting using PyTorch, OpenCV and DeepSORT, deployed on Streamlit.”fe8kv9Python · 63 ★“YOLOv9-FishEye: Improving method for fisheye camera object detection”WordPilotSvelte · 61 ★“WordPilot is a tool that converts YouTube videos into written blog format, with images and export options. It ”VayuAIPython · 60 ★“Vayuvahana Technologies Private Limited presents to you VajraV1, a state-of-the-art (SOTA) real time object de”PixEaglePython · 60 ★“Computer vision, object tracking, and target following for PX4 drones using OpenCV, YOLO, MAVSDK, MAVLink, and”CBPPython · 59 ★“Official Tensorflow Implementation of the AAAI-2020 paper "Temporally Grounding Language Queries in Videos by ”bilibili-analysis-helperPython · 59 ★“B站视频分析助手:提取字幕/评论/关键帧并生成深度分析报告。”OmniVideo-100KPython · 59 ★“OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains”Super-SloMo-tf2Python · 59 ★“Tensorflow 2 implementation of Super SloMo paper”watch-video-skillPython · 56 ★“A Claude skill that teaches your AI to watch videos. Use it to learn, absorb, copy, or give visual feedback li”bytetrack-pipPython · 56 ★“Packaged version of the ByteTrack repository”LVMAE-pytorchPython · 55 ★“Implementation of the proposed LVMAE, from the paper, Extending Video Masked Autoencoders to 128 frames, in Py”dji-tello-target-trackingPython · 55 ★“Modern autonomous drone tracking using YOLOv8 deep learning and PID control for DJI Tello drones with real-tim”OpenCV-Projects-cpp-pythonJupyter Notebook · 55 ★“Computer vision projects focused on object detection, object tracking, classical computer vision techniques, i”synthehiclePython · 54 ★“[WACVW 2023] A massive synthetic dataset for 3D multi-target multi-camera tracking and segmentation.”YOLOv8-object-tracking-blurring-countingJupyter Notebook · 53 ★“Real-Time Object Detection, Tracking, Blurring and Counting using YOLOv8”FramegenJavaScript · 53 ★“Real-time neural frame interpolation in the browser: hand-written WGSL runtime on raw WebGPU, 2.9 MB model, 2x”ThinkJEPAPython · 52 ★“ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model”aimotive_dataset52 ★“aiMotive public dataset”deep_vision_rosPython · 52 ★“ROS package for SOTA Computer Vision Models including SAM, Cutie, GroundingDINO, YOLO-World, VLPart, DEVA and ”small-vlm-sop-checkPython · 51 ★“Catch procedure errors before they become incidents. Train and evaluate small VLMs for temporal grounding in f”datagym-coreJava · 51 ★“Open source annotation and labeling tool for image and video assets”LTContextPython · 50 ★“[ICCV 2023] How Much Temporal Long-Term Context is Needed for Action Segmentation?”TESTAPython · 50 ★“[EMNLP 2023] TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding”YOLOv9-DeepSORT-Object-TrackingPython · 49 ★“YOLOv9 Object Tracking using PyTorch, OpenCV and DeepSORT”openfollowPython · 49 ★“OpenFollow enables you to track people and objects in 3D to automate lighting, audio and media. Generate PosiS”PointTADPython · 48 ★“[NeurIPS 2022] PointTAD: Multi-Label Temporal Action Detection with Learnable Query Points”PVT_ppPython · 48 ★“Full conference version of PVT++: A Simple End-to-End Latency-Aware Visual Tracking Framework”mcp-video-analyzerTypeScript · 47 ★“MCP server that turns any video — YouTube, Instagram, TikTok, Loom, X, Vimeo, direct URLs, local files — into ”yolo2647 ★“Ultralytics YOLO26 quickstart for detection, instance and semantic segmentation, depth estimation, classificat”hungarian_optimizerC++ · 47 ★“A C++ demo for Hungarian algorithm (Kuhn-Munkres algorithm).”BlobTracking.jlJulia · 46 ★“Detect and track blobs in video”video-helperPython · 45 ★“📺 AI视频学习助手: 自动生成 B站/YouTube/抖音/本地视频思维导图、笔记与总结。支持播客分析与视频索引,开源平替。AI Video Learning Assistant: Auto-generate Mind”ASM-LocPython · 45 ★“(CVPR2022) ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action Localization”multisensor-lmb-filtersMATLAB · 45 ★“Matlab implementations of various multi-sensor labelled multi-Bernoulli filters”vehicle_mtmcPython · 44 ★“Vehicle MTMC Tracking”mica-MovieCLIPPython · 43 ★“This repository contains the codebase for MovieCLIP: Visual Scene Recognition in Movies”nba_games43 ★“This dataset provides metadata, official statistics, and official play-by-play annotations for full-length NBA”VisionOSObjectTrackingDemoSwift · 43 ★“A demo project used for testing visionOS Object Tracking capabilities with Xcode, Reality Composer Pro, and Cr”Yolo8-ByteTrack-CSharpC# · 43 ★“ByteTrack implementation in C#, based on https://github.com/Vertical-Beach/ByteTrack-cpp”pixano-elementsTypeScript · 43 ★“Pixano Elements - Re-usable web components dedicated to data annotation tasks.”pixano-appJavaScript · 43 ★“Pixano App is a web-based smart-annotation tool for computer vision applications.”social-distancing-predictionPython · 42 ★“Out-of-the-box code and models for social distancing early forecasting.”frameflowTypeScript · 42 ★“FrameFlow is a high-fidelity multimedia platform that bridges the gap between raw video/image assets and creat”video_reason_benchPython · 41 ★“[ICLR 2026] "VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?", Yuanxin Liu, Kun Ou”TemporalBenchPython · 40 ★“TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models”YOLOv8-custom-object-detectionJupyter Notebook · 40 ★“This repository showcases the utilization of the YOLOv8 algorithm for custom object detection and demonstrates”