This project implements a Retrieval-Augmented Generation (RAG) system for financial documents, combining vector search with a large language model to provide accurate answers to financial queries.
| Category | Technologies |
|---|---|
| Language | |
| Vector DB | |
| NLP | |
| ML | |
| LLM | |
| Framework | |
| Data |
- Document processing and chunking
- Vector embedding generation.
- Efficient vector storage/retrieval
- Context-aware question answering
- Automatic model selection based on GPUs
- Prerequisites
- Python 3.9+
- GPU with sufficient VRAM (minimum 5GB recommended)
- Pinecone API key
- Hugging Face account (for Gemma model access)
git clone https://github.com/Shegun93/FinRAG.git
cd FinRAGpip install -r requirements.txt
Create a .env configuration files
PINECONE_API_KEY = "your-api-key"
PINECONE_ENVIRONMENT="region"huggingface-cli login
# Example query
query = "What was the operating profit increase from 2011-2012?"
answer = ask(query)
print(answer)
- Retrieve the most relevant document chunks
- Generate accurate answers using the Gemma LLM
- Display both the answer and the context used
- Chunk size: Adjust the chunk_size parameter in the split_into_chunks function
- Model selection: The system automatically selects the appropriate Gemma model based on available GPU memory
- Temperature: Control answer creativity via the temperature parameter in the ask function
- GPU Memory Errors: If you encounter memory issues, try:
- Using the 2B model instead of 7B
- Enabling
- Reducing the max_new_tokens parameter
- API key is correct
- Index name is unique
- Region matches your Pinecone configuration
This project is licensed under the MIT License
- Pinecone
- Huggingface
- Google for the Gemma models