Skip to content

Computer Vision

64 articles found

DeepSeek-AI Unveils Open-Source OCR Model with Human-Like Visual Processing Technology

DeepSeek-AI Unveils Open-Source OCR Model with Human-Like Visual Processing Technology

Jan 28, 2026
GitHub

DeepSeek-AI launches DeepSeek-OCR-2, an open-source visual OCR model featuring groundbreaking Visual Causal Flow technology that mimics human visual processing, supporting dynamic resolution up to 6×768×768 plus 1×1024×1024 image patches with document-to-markdown conversion, PDF processing, and streaming output capabilities through vLLM and Transformers frameworks.

Google Launches Gemini 3 Pro AI Model With Advanced Visual Reasoning and Document Processing Capabilities

Google Launches Gemini 3 Pro AI Model With Advanced Visual Reasoning and Document Processing Capabilities

Dec 08, 2025
Google

Google unveils Gemini 3 Pro, a breakthrough multimodal AI model that delivers state-of-the-art visual reasoning capabilities including complex document processing, pixel-precise spatial understanding, computer screen automation, and high-speed video analysis at 10+ FPS, promising major advances in education, medical imaging, and legal applications.

Previous
Page 4 of 7
Next
Showing 31 - 40 of 64 articles