Skip to content

About

Deep learning image captioning pipeline using CNN feature extraction and LSTM sequence generation, featuring an interactive web interface for real time inference.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Image Captioning Deep Learning System

An end-to-end deep learning system that analyzes visual input, extracts high-level semantic features using a Convolutional Neural Network (CNN), and generates natural language descriptive captions using a Recurrent Neural Network (LSTM).

Features

  • Automated visual feature extraction using a pre-trained CNN backbone
  • Sequence generation pipeline using LSTM networks trained with word tokenization
  • Greedy search and beam search decoding strategies for caption generation
  • Interactive web-based user interface for uploading custom images and previewing captions
  • Comprehensive Jupyter Notebook documenting data preprocessing, vocabulary building, and model training

System Architecture

Input Image
     ↓
CNN Feature Extractor (Encoder)
     ↓
Extracted Visual Features (Embedding)
     ↓
Tokenized Word Sequence (Decoder Context)
     ↓
LSTM Language Model
     ↓
Caption Prediction / Output Text

System Architecture

Project Structure

IMAGE-CAPTIONING-DEEP-LEARNING/
│
├── images/
│   ├── Block Diagram.png
│   ├── frontend.jpg
│   ├── Test1.png
│   ├── Test2.jpg
│   ├── DL_Proj_Image_1.jpg
│   ├── DL_Proj_Image_2.jpg
│   ├── DL_Proj_Image_3.jpg
│   └── DL_Proj_Image_4.jpg
│
├── notebook/
│   └── DL_Project.ipynb
│
├── .gitignore
├── README.md
└── requirements.txt

Demo & Results

Web Application Interface

Frontend Interface

Sample Model Predictions

Input Test Image Generated Caption Output
Test 1 Model-generated descriptive caption
Test 2 Model-generated descriptive caption

Setup Instructions

1. Clone the repository

git clone [https://github.com/umerharoon890/image-captioning-deep-learning.git](https://github.com/umerharoon890/image-captioning-deep-learning.git)
cd image-captioning-deep-learning

2. Create a virtual environment

python -m venv .venv

3. Activate the virtual environment

For Windows PowerShell:

.\.venv\Scripts\Activate.ps1

For macOS/Linux:

source .venv/bin/activate

4. Install dependencies

pip install -r requirements.txt

5. Run the Project

To inspect model training, evaluation, and tokenization:

jupyter notebook notebook/DL_Project.ipynb

To launch the web interface:

streamlit run app.py

Why This Project Is Useful

Bridging computer vision and natural language processing is fundamental for accessibility tools, automated visual documentation, and media indexing. This project demonstrates how multimodal deep learning architectures extract representations from convolutional layers and map them directly to natural language syntax.

About

Deep learning image captioning pipeline using CNN feature extraction and LSTM sequence generation, featuring an interactive web interface for real time inference.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages