An efficient Semantic Search System to Retrieve Photos using CLIP, BLIP and FAISS
DOI:
https://doi.org/10.51239/jictra.v16i1.357Keywords:
Multimodal Retrieval, Semantic Search, Personal Photo, Management, CLIP, BLIP, FAISS, Privacy PreservingAIAbstract
The rapid growth of personal photo collections has highlighted the limitations of traditional retrieval methods that rely on filenames or manual tagging, creating a persistent semantic gap between image content and user intent. This study proposes an efficient semantic search system that enables natural-language-based retrieval of personal photos through
a hybrid multimodal architecture. The system integrates CLIP for vision–language alignment, BLIP for automatic caption generation, and FAISS for high-speed similarity search, combining direct visual semantic matching with text-mediated understanding. Evaluated on a dataset of 2,000 personal images, the system demonstrates strong retrieval
performance, achieving mean cosine similarity scores of 0.254-0.282 and perfect Precision@6 (1.000) for attribute based queries. Results indicate high robustness to variations in lighting, pose, and background, with the strongest performance observed in clothing- and object-related searches, while event-based queries remain more challenging. Key contributions include a privacy-preserving on-device design, a dual-pathway multimodal workflow, and a comprehensive evaluation framework for semantic photo retrieval. Overall, the proposed system effectively bridges the semantic gap in personal photo management and demonstrates the practical value of multimodal AI for intuitive, human centred image retrieval.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Imtiaz Ali Dahri, Fida Hussain Dahri, Nisar Ahmed Dahri, Sajjad Hussain Bhutto

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.