Resume Search Engine and Finance Document QA
Two different ways to make a document collection useful through questions.
Built two retrieval systems for two different document problems. The first searched PDF resumes and returned relevant candidates with generated summaries. The second explored how a finance document system could check its evidence before answering.
The interesting part was not putting a chat box in front of a language model. It was deciding what to retrieve, how to keep the source document attached to the answer, and what to do when the first retrieved passage was not good enough.
Resume search
The resume application loaded PDF files, split their text into smaller pieces, created embeddings, and kept them in a persistent Chroma index. A natural language query retrieved the closest resumes, produced a summary for each result, kept page and source metadata, and exposed the results through a FastAPI service and a Streamlit interface. The original resume could still be opened from the result.
Finance documents
The finance project started with the familiar retrieve then answer flow and developed into a graph of retrieval and checking steps. It parsed documents, retrieved passages, judged document relevance, checked whether the answer was supported by the retrieved text, and tried retrieval again when the evidence was weak. Web search was a fallback rather than the first source of truth.
The difference between the two
The resume project is a practical search application with a clear result and source path. The finance project is a deeper experiment in making document question answering inspect its own work. Both are personal engineering projects, not claims of new retrieval research.
selected tools
- Python
- LangChain
- LangGraph
- Chroma
- Qdrant
- FastAPI
- Streamlit