Abstract

Modern retrieval systems must both be automated and interactive, allowing users to search and refine results in real time. We present AiSearch, a flexible multimodal retrieval framework that leverages the zero shot capabilities of Vision Language Models (VLMs) for natural language search over images and videos. AiSearch supports interactive search refinement through user feedback to tailor results to the user's intent, and allows visual benchmarking across multiple VLMs, enabling users to select the most suitable model for their task.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Koksal, A., Leong, M. C., Sintunata, V., Chin, C. L., & Fong, W. T. (2026). AiSearch: Interactive Multi-Modal Search with VLMs. https://omanscience.com/en/articles/aisearch-interactive-multi-modal-search-with-vlms

MLA 9

Koksal, Ali, et al. "AiSearch: Interactive Multi-Modal Search with VLMs." https://omanscience.com/en/articles/aisearch-interactive-multi-modal-search-with-vlms.

Chicago (author–date)

Koksal, Ali, Mei Chee Leong, Vicky Sintunata, Ching Ling Chin, and Wee Teck Fong. 2026. "AiSearch: Interactive Multi-Modal Search with VLMs." https://omanscience.com/en/articles/aisearch-interactive-multi-modal-search-with-vlms.

Harvard

Koksal, A., Leong, M. C., Sintunata, V., Chin, C. L. and Fong, W. T. (2026) 'AiSearch: Interactive Multi-Modal Search with VLMs', Available at: https://omanscience.com/en/articles/aisearch-interactive-multi-modal-search-with-vlms.

Vancouver

Koksal A, Leong MC, Sintunata V, Chin CL, Fong WT. AiSearch: Interactive Multi-Modal Search with VLMs. https://omanscience.com/en/articles/aisearch-interactive-multi-modal-search-with-vlms

IEEE

A. Koksal, M. C. Leong, V. Sintunata, C. L. Chin, and W. T. Fong, "AiSearch: Interactive Multi-Modal Search with VLMs," https://omanscience.com/en/articles/aisearch-interactive-multi-modal-search-with-vlms.