Ask questions about your own documents — answers are grounded in your text, cited to the source, and it says “I don't know” instead of guessing.
Chunking splits documents into overlapping windows so context isn't cut at boundaries. Embeddings here use a TF-IDF vectoriser computed in your browser; retrieval ranks chunks by cosine similarity. Answers are built only from retrieved chunks and cited; below a similarity threshold the assistant returns “I don't know” rather than hallucinating. This static demo is extractive (returns the most relevant passage). The full Python version adds neural sentence-transformers embeddings and optional LLM generation. Uploaded files are processed in-memory and never leave your device. 🔗 Source on GitHub