Overview

Find is a local-first image intelligence platform: it uploads, indexes, searches and clusters images on your own machine. Image processing, vector generation and search all stay inside your local stack, so a personal photo library never has to leave it.

  • Upload individual images or ZIP archives.
  • Extract captions, detected objects, OCR text, EXIF metadata and dimensions.
  • Search with natural language over hybrid embeddings.
  • Cluster related images automatically once indexing completes.
  • Share albums with scoped links, and keep hidden images in a password-gated vault.

Architecture

Find's architecture: a Next.js frontend, a FastAPI backend with PostgreSQL, pgvector, Redis, RQ workers and MinIO, and the local ML pipeline

The frontend is Next.js with React Query. The backend is FastAPI with SQLAlchemy, PostgreSQL and pgvector, Redis with RQ workers, and MinIO for object storage. The ML pipeline runs YOLO for detection, BLIP for captions, PaddleOCR for text, SigLIP for embeddings, InsightFace for faces and HDBSCAN for clustering.

It ships as Docker Compose profiles, so the same app runs with no AI at all, with deterministic mock outputs for tests, on CPU, or on an NVIDIA GPU.

Results

images processed
100k+
recognition accuracy
98%
query latency
−350 ms
index size
−20%
Pipeline after optimization, baseline = 100
Pipeline after optimization, baseline = 100
MetricBeforeAfterChange
Image throughput100130+30%
Memory usage10085−15%
Index size10080−20%

Vector quantization and compression cut the index by 20%, which is what made 100,000+ images practical on one machine. Resolving three pipeline bottlenecks took 350 ms off query latency.

Screens

Find's natural-language search page
Search
Automatically generated image clusters
Clusters
Uploading images and ZIP archives
Upload

Open source

Find is AGPL-3.0 licensed and open to contributors through GirlScript Summer of Code 2026, with issues labelled by difficulty so first-time contributors know where to start.