🔥[CVPR2025] EventGPT: Event Stream Understanding with Multimodal Large Language Models
-
Updated
Jul 26, 2025 - Python
🔥[CVPR2025] EventGPT: Event Stream Understanding with Multimodal Large Language Models
Multimodal Brain Tumor Segmentation Challenge 2018
2022京东全球人工智能技术创新大赛 电商关键属性的图文匹配任务第1名方案
An agentic, zero‑shot document intelligence engine that sees, understands, and extracts from any PDF, no training, no hallucinations. Just define your fields and get trusted, structured outputs with confidence scores, deployed locally and built for the enterprise.
Developed an end-to-end Multimodal Retrieval-Augmented Generation (RAG) ingestion pipeline for processing complex PDF documents containing text, tables, and images. • Implemented high-resolution PDF parsing using Unstructured.io with OCR support for scanned documents, enabling extraction text, tables, and embedded images while preserving document
To associate your repository with the mutilmodel topic, visit your repo's landing page and select "manage topics."