Research & talks
Ideas worth sharing.
Work on multimodal retrieval, video-aware agents, computer vision, and scalable ML systems.
Selected papers
2026
Tool-Skills vs. Workflow-Skills: Routing and Composition in a Multimodal Agent Harness
Case study of Tinycloud's multimodal video agent harness, showing how tool-skills and workflow-skills affect routing and composition.
Related work at Cloudglue2025
Smart Routing for Multimodal Video Retrieval: When to Search What
Routes queries across ASR, OCR, and visual indices to reduce retrieval cost while preserving effectiveness.
Related work at Cloudglue2025
VideoMCP: Video Intelligence for Multimodal Agent Reasoning
Connects agents to speech, OCR, and visual analysis through MCP, with an evaluation against a speech-only baseline on enterprise video.
Related work at Cloudglue2025
RAVEN: An Agentic Framework for Multimodal Entity Discovery from Large-Scale Video Collections
Extracts structured entities from large video collections using adaptive schemas and multimodal context.
Autonomous Video Hunter project2025
Video Enriched Retrieval Augmented Generation Using Aligned Video Captions
Uses aligned visual captions to make video evidence more usable inside retrieval-augmented assistants.
Related work at Cloudglue2023
FORB: A Flat Object Retrieval Benchmark for Universal Image Embedding
Introduces a benchmark for evaluating universal image embeddings on flat-object visual retrieval, including logos and diverse 2D patterns.
From the benchmark to mDexTalks & demonstrations
2026
Evidence-Grounded Video Investigation over Large Video Corpora
Hands-on demonstration of answering questions over large video corpora with rich visualizations.
Overcast project & code2026
Overcast: Video OSINT Agent. Point It at 100 Videos, Ask Anything
A video OSINT agent that ingests large video collections and answers investigative questions on demand.
Overcast project & code2025
Autonomous Video Hunter: AI Agents for Real-Time OSINT
A research-to-demo bridge for video-aware agents operating over live investigative workflows.
Autonomous Video Hunter project2025
AI That Gets the Picture: Building Video-Aware Assistants
Model Context Protocol, multimodal RAG, and TypeScript patterns for assistants that understand video.
Related work at Cloudglue2020
MLOps at Snapchat: Continuous Machine Learning with Kubeflow and Spinnaker
Production ML operations for training, evaluating, and deploying perception systems.
ML infrastructure at Snap2019
Build Accurate Training Datasets with Amazon SageMaker Ground Truth
Data-labeling and training-pipeline lessons from building computer vision systems at Snap.
ML infrastructure at SnapPlanning a talk?
I speak about video agents, multimodal retrieval, and building production ML systems.