This project is a Streamlit web application that leverages OpenAI's GPT-4o to generate descriptions for uploaded images
-
Updated
Jul 19, 2024 - Python
This project is a Streamlit web application that leverages OpenAI's GPT-4o to generate descriptions for uploaded images
AI Image Description Generator accurately extracts the key elements from images and interprets the creative purposes behind them, which can be applied in fields such as scientific research, artistic creation, and the mutual search between images and texts.
Dsh-visual-plugin.Give your text-only model eyes: forward user images to any OpenAI-compatible vision model and see the results in a Web UI right panel
Image identification with Kosmos2 model, drawing and cutting bbox with object detection
为 DeepSeek V4.0 等单模态大模型装上眼睛和耳朵 —— 本地视觉+音频 MCP 服务器,基于 Ollama + MiniCPM-V 4.6 + faster-whisper,图片描述·视频分析·语音转文字
Experimenting with mastodon.social client alt-text usage dataset.
DeepSeek Harness 识图插件:为不具备原生识图能力的模型提供识图能力(阿里云百炼 qwen3.5-omni-plus,失败自动切换智谱 glm-4.6v-flash)。由 claude-vision-skill 移植适配。 | Vision tool for DeepSeek Harness
Give vision to non-vision models in OpenCode using Google's free Gemini API. Transparently replaces image parts with detailed text descriptions — vision models stay untouched.
279K image alt-text pairs from 489 Bluesky accounts — curated for quality, validated at 90%+ alt-text rate
A DeepSeek Harness tool plugin that lets text-only agents "see" local images — auto-detects the real format and returns a detailed text description via any OpenAI-compatible vision model.
A lightweight console utility that uses an LLM to generate descriptions and keywords for images.
A new package that processes user-submitted text descriptions of images or videos containing watermarks and returns structured, watermark-free descriptions. It uses an LLM to reinterpret the content w
给 DeepSeek Harness 纯文本模型装上原生视觉(Windows):粘贴即看图——预注入描述,模型首轮就看见,不用选模型、不用调工具;see_image 精查;自定义视觉后端(任意 OpenAI 兼容模型)+ 四后端容灾;换主模型视觉自动跟随。| Give text-only DeepSeek Harness models native-feeling vision on Windows: paste and the model just sees it — pre-injected descriptions, see_image tool, custom backends, 4-backend failover.
An intelligent assistant powered by the ReAct framework, leveraging LangChain for tool-based reasoning and Gradio for a user-friendly interface. Supports tasks like weather queries, PDF summarization, image descriptions, and more.
A see_image vision tool plugin for DeepSeek Harness — describe images through any OpenAI-compatible vision model (GitHub Copilot, OpenAI, Ollama, vLLM, LM Studio).
📊 Multi-Modal RAG 2.0 — An enterprise-grade RAG system that processes PDFs, tables, charts/images, and audio files together. Uses Unstructured.io, Camelot, Gemini Vision, Whisper, ChromaDB, FastAPI, and Streamlit. Increases information retrieval accuracy by 35% over text-only RAG.
It is an innovative repository housing a sophisticated Large Language Model (LLM) project, showcasing the intersection of advanced natural language processing and cutting-edge artificial intelligence. This repository serves as a comprehensive platform for the development, experimentation, and application of state-of-the-art language models.
AI-Powered-Solution-for-Assisting-Visually-Impaired-Individuals
CLI tool for generating image metadata in bulk — powered by Gemini for smarter, more context-aware descriptions.
DSH multimodal plugin: drag-image auto-describe with a configurable OpenAI-compatible vision model (host patch + agent preset + optional adapter). MIT.
To associate your repository with the image-description topic, visit your repo's landing page and select "manage topics."