Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

78 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

VidChain

High-Fidelity Multimodal RAG Framework for Forensic Video Intelligence

PyPI Engine RAG Vector DB Orchestration Local First

Python CUDA Downloads PyPI version License GitHub stars

VidChain Banner

VidChain is a local-first multimodal RAG framework powered by the IRIS Engine (Intelligent Retrieval & Insight System). It decomposes video into visual, auditory, OCR, and temporal signal streams and fuses them into a queryable intelligence layer, intended for forensic analysis, security auditing, and automated video summarization with on-device privacy by default.


Table of Contents

  1. Overview
  2. Features
  3. Installation
  4. Configuration
  5. Quick Start
  6. CLI Reference
  7. SDK: Modular Sensor Matrix
  8. REST API
  9. Architecture
  10. Troubleshooting
  11. Contributing
  12. License

Overview

VidChain turns raw video into a queryable intelligence layer. Each ingested video is processed through a modular sensor pipeline (visual, audio, OCR, motion, behavioral), decomposed into an isolated Temporal Knowledge Graph, and fused with vector retrieval to produce grounded, timestamp-cited answers. Inference runs entirely on-device by default; cloud models are supported as an explicit opt-in, not a requirement.


Features

Capability Description
4-Route Agentic Router Classifies queries into Narrative Summarization, Local Forensic Search, Global Master Intelligence, and Conversational Dialogue
Global Master Intelligence Cross-video entity tracking via a macro-graph, enabling pattern recognition across isolated sessions
Temporal Persistence Chronological reasoning that bridges frame gaps and maintains state continuity between sensor logs
Recursive Map-Reduce Summarizer Collapses hours of video into coherent reports without hitting LLM context limits
Neural Concurrency Locking Prevents state corruption during simultaneous ingestion and query operations
Local-First Execution Vision (Moondream) and reasoning (Llama 3 via Ollama) run on-device by default; no data leaves the machine unless a cloud model is explicitly configured

Installation

Prerequisites

Requirement Version
Python 3.11+
CUDA 12.1+
Ollama Latest (running)
Node.js v18+ (for web portal)

Option A — Install from PyPI (recommended)

pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
pip install vidchain

Option B — Install from source (for development)

pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
git clone https://github.com/rahulsiiitm/videochain-python
cd videochain-python
pip install -e .

Model weights (required either way)

ollama pull moondream   # Vision Language Model
ollama pull llama3      # Language Model for reasoning & routing

CPU Fallback: If no CUDA device is detected, VidChain automatically degrades to CPU mode — no code changes required.


Configuration

Variable Required Description
GEMINI_API_KEY Only if using a Gemini model via --llm gemini/... API key for Google's Gemini models, used through LiteLLM
db_path (constructor arg, not env var) Yes Local directory where ChromaDB vectors and Temporal Knowledge Graphs are stored

No other environment variables are required for the default local-only configuration. Any LiteLLM-compatible provider can be substituted for --llm or --vlm; check the LiteLLM provider docs for the corresponding key name if using a provider other than Gemini or Ollama.


Quick Start

from vidchain import VidChain

vc = VidChain(db_path="./forensic_vault")

# Ingest video (runs full default pipeline)
video_id = vc.ingest(video_source="interview_01.mp4")

# Query
response = vc.ask("What is the main topic of discussion?", video_id=video_id)
print(response["text"])

# Summarize
summary = vc.summarize_video(video_id=video_id, mode="concise")
print(summary)

CLI Reference

vidchain-serve

Launches the FastAPI backend and Next.js dashboard.

vidchain-serve
  • API available at http://localhost:8000
  • Dashboard opens at http://localhost:3000
  • Includes a 7-second warmup before accepting requests

vidchain-analyze

Headless video ingestion from the terminal.

vidchain-analyze path/to/video.mp4 --vlm moondream
Flag Description
--vlm <model> Vision model to use (default: moondream, local)
--llm <model> Reasoning model to use (default: ollama/llama3, local)
--fast Replaces VLM with YOLO for high-speed detection (ideal for long CCTV footage)
--emotion Injects DeepFace emotion analysis node
--action Injects MobileNetV3 action classification node

Model substitution: VidChain uses LiteLLM, so any compatible model can be swapped in, including cloud models if higher reasoning quality is preferred over on-device execution:

# Local (default)
vidchain-analyze video.mp4 --llm "ollama/llama3"

# Cloud (opt-in, requires API key export)
export GEMINI_API_KEY="your_api_key"
vidchain-analyze video.mp4 --llm "gemini/gemini-2.5-flash"

# Custom VLM
vidchain-analyze video.mp4 --vlm "llava:7b"

SDK: Modular Sensor Matrix

VidChain uses a LangChain-inspired composable pipeline. Each Node handles one sensing modality; chains are assembled per use case.

Available Nodes

Node Modality Description
AdaptiveKeyframeNode Logic Gaussian-differential sampling — drops redundant frames to reduce compute load
LlavaNode Visual Scene semantics, descriptive captions, and situational context
YoloNode Visual High-speed discrete object detection (lightweight fallback for LlavaNode)
WhisperNode Audio Speech transcription and acoustic anomaly detection (e.g., shouts)
OcrNode Text Digital trace extraction — license plates, screens, documents
TrackerNode Motion Persistent object tracking (IoU) and camera motion estimation (Optical Flow)
EmotionNode Behavioral Facial sentiment analysis
ActionNode Behavioral Human activity classification via MobileNetV3

Custom Pipeline Example

from vidchain import VidChain
from vidchain.pipeline import VideoChain
from vidchain.nodes import AdaptiveKeyframeNode, LlavaNode, OcrNode, TrackerNode

vc = VidChain(db_path="./forensic_vault")

surveillance_chain = VideoChain(nodes=[
    AdaptiveKeyframeNode(change_threshold=1.5),  # High sensitivity
    LlavaNode(model="moondream"),
    OcrNode(),
    TrackerNode()
])

video_id = vc.ingest(
    video_source="gate_camera_04.mp4",
    chain=surveillance_chain
)

response = vc.ask(
    "Were there any vehicles with visible license plates after 14:00?",
    video_id=video_id
)
print(response)

REST API

Exposed when running vidchain-serve.

Method Endpoint Description
GET /api/health System status and list of ingested video IDs
POST /api/sessions Create a new isolated neural session
POST /api/ingest Submit a video file path for background processing
POST /api/query Run a natural language query through the Agentic Router
GET /api/media-stream Serve local video securely for frontend playback

Architecture

Isolated GraphRAG

Each ingested video generates a dedicated Temporal Knowledge Graph (.pkl). The RAG engine retrieves semantically relevant chunks from ChromaDB and fuses them with structured graph data (co-occurrences, tracking IDs, timestamps). Memory boundaries are strictly enforced — no cross-video context bleed.

The Neural Lens

Every query response is paired with a Base64-encoded visual snapshot extracted directly from the referenced timestamp, providing visual grounding for AI-generated claims.


Troubleshooting

Symptom Likely Cause Fix
vidchain-serve fails to start Ollama not running Start Ollama before launching VidChain (ollama serve)
Ingestion runs but very slowly No CUDA device detected, running on CPU fallback Confirm nvidia-smi shows a GPU; reinstall the CUDA-enabled Torch build from Installation
--llm gemini/... fails with an auth error GEMINI_API_KEY not exported export GEMINI_API_KEY="your_api_key" before running the command
Dashboard loads but shows no videos Wrong db_path between ingest and query calls Ensure VidChain(db_path=...) points to the same directory across sessions

Contributing

Issues and pull requests are welcome via GitHub Issues. For substantial changes, open an issue first to discuss scope before submitting a PR.


License

MIT — See LICENSE for details.

Author: Rahul Sharma — IIIT Manipur Portfolio · GitHub

Star this repo if you find it useful — it helps the IRIS Engine grow.

Releases

Packages

Contributors

Languages