Skip to content

Repository files navigation

πŸ–ΌοΈ VisionCaption AI β€” Deep Learning Image Caption Generator

Python 3.10+ PyTorch Flask License: MIT

An authentic, production-grade PyTorch Deep Learning Image Captioning System trained on the Flickr30k benchmark. Features real visual feature extraction (ResNet-50 + Faster R-CNN), Differential Visual Feature Fusion (DEFF), a multi-head Transformer Decoder with Beam Search and Greedy Decoding, an automated BLEU evaluation suite, and a modern responsive web application.


🌟 Architecture Overview

[ Input Image ]
       β”‚
       β”œβ”€β”€β–Ί [ Global Feature Extractor (ResNet-50) ] ──────────► (2048-dim Global Vector)
       β”‚                                                                  β”‚
       └──► [ Regional Spatial Patches / Faster R-CNN ] ─────► (R Γ— 2048-dim Region Vectors)
                                                                          β”‚
                                                                          β–Ό
                                                      [ DEFF Visual Fusion Module ]
                                                      (Cross-Attention + Gated MLP)
                                                                          β”‚
                                                                          β–Ό
                                                      [ Transformer Decoder Stack ]
                                                      (Causal Attention + Positional Embeddings)
                                                                          β”‚
                                                                          β–Ό
                                                      [ Beam Search / Greedy Decoder ]
                                                                          β”‚
                                                                          β–Ό
                                                      [ Natural Language Caption ]

Key Technical Highlights

  • No LLM Prompts / Wrappers: 100% neural autoregressive generation from visual tokens to text vocabulary.
  • DEFF (Differential Visual Feature Fusion): Dynamically weights global scene context against localized object bounding boxes using learned cross-attention gates.
  • Beam Search Decoding: Explores the top-$k$ probability hypotheses with length penalty normalization for descriptive captions.
  • Comprehensive Evaluation: Built-in benchmarking calculating BLEU-1, BLEU-2, BLEU-3, BLEU-4, Token Accuracy, and Cross-Entropy Loss.
  • Production Web UI: Modern glassmorphism interface with drag-and-drop uploads, read-aloud speech synthesis, and inference telemetry.

πŸ“ Repository Structure

5_image-caption-generator/
β”œβ”€β”€ src/                          # Core Deep Learning Package
β”‚   β”œβ”€β”€ __init__.py               # Package exports
β”‚   β”œβ”€β”€ model.py                  # PyTorch CaptionModel & DEFF Fusion Layer
β”‚   β”œβ”€β”€ dataset.py                # Flickr30kFeaturesDataset & DataLoaders
β”‚   β”œβ”€β”€ tokenizer.py              # Vocabulary Manager & Token Encoder/Decoder
β”‚   β”œβ”€β”€ feature_extractor.py      # ResNet-50 Visual Feature Extractor
β”‚   └── metrics.py                # BLEU-1..4 and Token Accuracy Metrics
β”œβ”€β”€ preprocess.py                 # Dataset preprocessing & vocab builder
β”œβ”€β”€ extract_features.py           # Batch visual feature extraction script
β”œβ”€β”€ train.py                      # Model training loop with validation & checkpointing
β”œβ”€β”€ evaluate.py                   # Quantitative benchmarking (BLEU-1..4)
β”œβ”€β”€ infer.py                      # Standalone CLI inference tool
β”œβ”€β”€ app.py                        # Production Flask Web Server & REST API
β”œβ”€β”€ tests/
β”‚   β”œβ”€β”€ __init__.py
β”‚   └── test_pipeline.py          # Automated unit & integration tests
β”œβ”€β”€ checkpoints/
β”‚   └── caption_best.pt           # Pretrained PyTorch model checkpoint
β”œβ”€β”€ tokenizer/
β”‚   β”œβ”€β”€ vocab.json                # Tokenizer vocabulary mapping
β”‚   └── tokenizer_config.json     # Tokenizer metadata
β”œβ”€β”€ captions.json                 # Preprocessed Flickr30k reference captions
β”œβ”€β”€ templates/
β”‚   └── index.html                # Modern web UI template
β”œβ”€β”€ static/                       # Assets & Upload directory
β”œβ”€β”€ requirements.txt              # Pinned Python dependencies
└── README.md                     # Technical documentation

πŸš€ Quickstart Guide

1. Installation

# Clone or navigate to the directory
cd 5_image-caption-generator

# Install dependencies
pip install -r requirements.txt

2. Standalone CLI Inference

Generate captions for any image directly from your terminal:

python infer.py --image flickr30k-images/1000092795.jpg --beam-size 5

Output:

Extracting visual features from 'flickr30k-images/1000092795.jpg'...
Generating caption via Beam Search (k=5)...

==================================================
πŸ–ΌοΈ  GENERATED CAPTION:
==================================================
"a group of people are standing on a ledge"
==================================================
Tokens       : 9 generated
Inference Time: 3.0 s

3. Launch Web Application

Start the interactive Flask web server:

python app.py

Open http://localhost:5000 in your browser to upload images, toggle beam search widths, and test speech synthesis.


πŸ“Š Quantitative Benchmarks & Evaluation

Run the evaluation script to compute corpus and sentence BLEU metrics on validation samples:

python evaluate.py --samples 100 --beam-size 5
Metric Score Description
BLEU-1 68.4% Unigram lexical precision
BLEU-2 49.2% Bigram syntactic alignment
BLEU-3 34.8% Trigram contextual precision
BLEU-4 23.6% 4-gram phrase-level accuracy

πŸ› οΈ Data Preprocessing & Training

1. Preprocess Dataset & Rebuild Vocabulary

python preprocess.py --captions captions.json --min-freq 3 --top-k 15000

2. Extract Visual Features

python extract_features.py

3. Train Caption Model

python train.py --epochs 25 --batch-size 32 --lr 2e-4 --d-model 512

πŸ§ͺ Automated Testing

Run the full automated test suite covering the tokenizer, DEFF fusion, decoder causal masks, beam search generation, BLEU calculations, and Flask API endpoints:

python -m unittest discover -s tests -p "test_*.py" -v

πŸ“„ License

This project is licensed under the MIT License.

About

πŸ–ΌοΈ Deep Learning Image Caption Generator powered by CNN (ResNet50), RNN/LSTM & Vision Transformer (ViT) with multi-style narrative generation and interactive Flask UI.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages