Medical Chatbot — Retrieval-augmented Q&A
A Flask, Pinecone, and LangChain retrieval app with a containerized deployment path to AWS EC2.
- Role
- Python / AI Developer
- Timeline
- 2024
- Status
- Personal project
- Area
- Retrieval / RAG
Overview
A retrieval-augmented medical assistant that surfaces relevant health information through semantic search and context-aware generation. The focus was turning a knowledge base into a usable chat interface, with a containerized deployment path to AWS.
Problem
People want quick, understandable medical information, but in a high-trust domain answers need to stay tied to source material.
Approach
Semantic retrieval combined with LLM generation, so answers are informed by retrieved domain content.
01
Index
Embeddings → Pinecone
02
Retrieve
LangChain
03
Generate
OpenAI GPT
04
Serve
Flask, Docker, EC2
A Flask app serves the chat, Pinecone stores the vector index, LangChain orchestrates retrieval, and OpenAI generates answers over the retrieved context.
Outcomes
- Built the full workflow from embedding generation to chat interface.
- Used Pinecone retrieval so answers are based on indexed source content.
- Explored containerized deployment and CI/CD-oriented hosting on AWS.
Challenges
- Balancing usefulness with the sensitivity of the domain.
- Keeping retrieval quality high across varied medical questions.
- Reducing latency without oversimplifying the pipeline.
Stack
Python · Flask · LangChain · Pinecone · OpenAI GPT · Docker · AWS EC2