AI-powered English pronunciation analysis application using Google Gemini 2.0 Flash.
- 🎤 Pronunciation Error Detection - Identify mispronounced or omitted words
- 📊 IELTS Speaking Metrics - Evaluate fluency, lexical resource, grammar, and pronunciation
- 📝 Speaking Reports - Generate comprehensive performance reports
- 🎵 Audio Support - MP3, WAV, WebM formats (max 16MB)
- 🧹 Auto Cleanup - Automatic deletion of old uploads (7-day retention)
- 🛡️ Rate Limiting - IP-based limits to prevent spam and control API costs
- Backend: Flask + Python 3.11
- AI: Google Gemini 2.0 Flash via LangChain
- Server: Gunicorn + Nginx (production)
- Deployment: AWS EC2, Docker, CloudFormation
- Google API Key - Get from Google AI Studio
- Python 3.11+ or Docker
# 1. Clone repository
git clone <your-repo-url>
cd pronuciation_checker
# 2. Create virtual environment
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
# 3. Install dependencies
pip install -r requirements.txt
# 4. Configure environment
cp .env.example .env
nano .env # Add your GOOGLE_API_KEY
# 5. Run application
python run.py
# 6. Test
curl http://localhost:5000/api/v1/health-check# 1. Configure environment
cp .env.example .env
nano .env # Add your GOOGLE_API_KEY
# 2. Deploy
bash scripts/deploy_docker.sh
# 3. Access
curl http://localhost:5000/api/v1/health-check# 1. Setup infrastructure
bash scripts/setup_ec2.sh
# 2. SSH to instance
ssh -i pronunciation-checker-key.pem ubuntu@YOUR_IP
# 3. Upload code and deploy
bash scripts/deploy_app.sh
# 4. Configure API key
nano /home/ubuntu/pronuciation_checker/.env
# 5. Restart
sudo systemctl restart pronunciation-checkerbash scripts/deploy_cloudformation.shCost: ~$18-22/month (FREE with AWS Free Tier for 12 months)
Rate Limit: 10 requests per hour per IP
POST /api/v1/analyze-pronunciation-error
Content-Type: multipart/form-data
Parameters:
- audio: file (mp3/wav/webm)
- text: string (reference text)
Response:
{
"status": "success",
"data": {
"errors": [
{
"word": "example",
"position": 0,
"error_type": "phát âm sai",
"correct_pronunciation": "/ɪɡˈzæmpəl/",
"your_pronunciation": "/ɪɡˈzɑːmpəl/",
"explanation": "Nguyên âm sai..."
}
],
"html_output": "<span>...</span>"
}
}Rate Limit: 10 requests per hour per IP
POST /api/v1/evaluate-speech-metrics
Content-Type: multipart/form-data
Parameters:
- audio: file
- text: string
Response:
{
"status": "success",
"data": {
"fluency_and_coherence": {
"score": 7,
"feedback": "..."
},
"lexical_resource": {
"score": 6,
"feedback": "..."
},
"grammatical_range_and_accuracy": {
"score": 7,
"feedback": "..."
},
"pronunciation": {
"score": 6,
"feedback": "..."
}
}
}Rate Limit: 20 requests per hour per IP
POST /api/v1/generate-speaking-report
Content-Type: application/x-www-form-urlencoded
Parameters:
- text: string (test results)
Response:
{
"status": "success",
"data": {
"overall_assessment": "...",
"common_errors": [...],
"improvement_suggestions": [...]
}
}Rate Limit: 100 requests per hour per IP
# Get storage statistics
GET /api/v1/storage-stats
# Manual cleanup
POST /api/v1/cleanup-uploads
Content-Type: application/json
{"max_age_days": 7}
# Health check
GET /api/v1/health-checkCreate a .env file:
# Required
GOOGLE_API_KEY=your_google_api_key_here
# Optional
SECRET_KEY=your-secret-key
UPLOAD_FOLDER=./uploads
# Cleanup (defaults shown)
CLEANUP_ENABLED=true
CLEANUP_MAX_AGE_DAYS=7
CLEANUP_INTERVAL_HOURS=24
# Rate Limiting (defaults shown)
RATELIMIT_ENABLED=true
RATELIMIT_AI_ENDPOINTS=10 per hour
RATELIMIT_REPORT_ENDPOINT=20 per hour
RATELIMIT_UTILITY_ENDPOINTS=100 per hourTo prevent API abuse and control Google API costs, the application implements IP-based rate limiting. This is especially important for demo deployments where API credits are limited.
| Endpoint Category | Limit | Endpoints |
|---|---|---|
| AI Operations | 10 requests/hour | /analyze-pronunciation-error, /evaluate-speech-metrics |
| Report Generation | 20 requests/hour | /generate-speaking-report |
| Utility | 100 requests/hour | /health-check, /storage-stats, /cleanup-uploads |
All responses include rate limit information in headers:
X-RateLimit-Limit: 10 # Maximum requests allowed
X-RateLimit-Remaining: 7 # Requests remaining in current window
X-RateLimit-Reset: 1640995200 # Unix timestamp when limit resets# View rate limit headers
curl -I http://localhost:5000/api/v1/health-check
# Example response headers:
# X-RateLimit-Limit: 100
# X-RateLimit-Remaining: 99
# X-RateLimit-Reset: 1640995200{
"error": "Rate limit exceeded",
"message": "Too many requests. This is a demo application with limited API credits. Please try again later.",
"limit": "10 per 1 hour"
}HTTP Status Code: 429 Too Many Requests
Edit .env file:
# Disable rate limiting (not recommended for production)
RATELIMIT_ENABLED=false
# Adjust limits (examples)
RATELIMIT_AI_ENDPOINTS=5 per hour # Stricter
RATELIMIT_AI_ENDPOINTS=20 per hour # More generous
RATELIMIT_AI_ENDPOINTS=100 per day # Daily limit
RATELIMIT_REPORT_ENDPOINT=50 per hour
RATELIMIT_UTILITY_ENDPOINTS=200 per hourProblem: Getting 429 errors too frequently
Solutions:
- Increase limits in
.envfile - Wait for the rate limit window to reset
- Check
X-RateLimit-Resetheader for reset time - For production use, consider implementing user authentication with per-user limits
Problem: Rate limiting not working
Solutions:
- Verify
RATELIMIT_ENABLED=truein.env - Check application logs for rate limiter initialization
- Ensure
flask-limiteris installed:pip install flask-limiter
Client Upload → Validate → Read to RAM → Base64 → Gemini API → Response → Discard
Files are NOT saved to disk by default for:
- 🚀 Better performance (no disk I/O)
- 🔒 Enhanced privacy (files don't persist)
- 💾 Zero storage usage
Uncomment line 33 in app/routes/api.py:
# Change from:
# audio_path = save_uploaded_file(audio_file)
# To:
audio_path = save_uploaded_file(audio_file)Then files will be saved to ./uploads/ and automatically cleaned up after 7 days.
- 🤖 Automatic - Runs every 24 hours
- 🗑️ Smart - Deletes files older than 7 days
- 📊 Monitored - API endpoints for stats
- 🔄 Redundant - Background scheduler + cron job
- App starts → Cleanup runs immediately
- Every 24 hours → Cleanup runs automatically
- Files older than 7 days → Deleted
- Empty directories → Removed
- All operations → Logged
CLEANUP_ENABLED=true # Enable/disable
CLEANUP_MAX_AGE_DAYS=7 # Delete after X days
CLEANUP_INTERVAL_HOURS=24 # Run every X hours# View logs
sudo journalctl -u pronunciation-checker | grep -i cleanup
# Check storage
curl http://localhost:5000/api/v1/storage-stats
# Manual cleanup
curl -X POST http://localhost:5000/api/v1/cleanup-uploadspronuciation_checker/
├── app/
│ ├── AI_module/ # AI workflows and nodes
│ ├── routes/ # API endpoints
│ ├── services/ # Business logic & scheduler
│ └── utils/ # Utilities & cleanup
├── scripts/
│ ├── setup_ec2.sh # AWS infrastructure setup
│ ├── deploy_app.sh # Application deployment
│ ├── deploy_docker.sh # Docker deployment
│ ├── deploy_cloudformation.sh
│ └── cleanup_cron.sh # Cron job for cleanup
├── cloudformation/
│ └── pronunciation-checker.yaml
├── Dockerfile
├── docker-compose.yml
├── requirements.txt
├── run.py
└── .env # Create this from .env.example
# Run application
python run.py
# Install dependencies
pip install -r requirements.txt
# Run tests
pytest tests/# View logs
docker-compose logs -f
# Restart
docker-compose restart
# Stop
docker-compose down
# Rebuild
docker-compose up -d --build# View logs
sudo journalctl -u pronunciation-checker -f
# Restart service
sudo systemctl restart pronunciation-checker
# Check status
sudo systemctl status pronunciation-checker
# Update code
cd /home/ubuntu/pronuciation_checker
git pull
sudo systemctl restart pronunciation-checker# Check logs
sudo journalctl -u pronunciation-checker -n 100
# Verify dependencies
source venv/bin/activate
pip list# Find process
sudo lsof -i :5000
# Kill process
sudo kill -9 PID- Verify API key in
.env - Check quota in Google Cloud Console
- Ensure Gemini API is enabled
# Check memory
free -h
# Add swap
sudo fallocate -l 2G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile- ✅ Never commit
.envfile - ✅ Use strong
SECRET_KEY - ✅ Restrict SSH access (security groups)
- ✅ Enable HTTPS with Let's Encrypt
- ✅ Regular updates:
sudo apt update && sudo apt upgrade - ✅ Monitor CloudWatch for unusual activity
- ✅ Use IAM roles instead of access keys
- Max file size: 16MB
- Supported formats: MP3, WAV, WebM
- Processing time: 2-5 seconds (depends on audio length)
- Concurrent requests: Handled by Gunicorn workers
- Memory usage: ~200-500MB per worker
- Instance: ~$15/month
- Storage (20GB): ~$2/month
- Data transfer: ~$1-5/month
- Total: ~$18-22/month
- t2.micro instance FREE for 12 months
- 30GB storage FREE for 12 months
- Perfect for testing!
- Create route in
app/routes/api.py - Add service logic in
app/services/ - Update AI workflows in
app/AI_module/if needed - Test locally
- Deploy
Edit prompts in app/AI_module/nodes.py:
analyze_pronunciation_errors_nodeevaluate_speech_metrics_nodegenerate_speaking_report_node
[Your License Here]
[Your Contributing Guidelines Here]
For issues and questions:
- Check logs first
- Review error messages
- Verify environment variables
- Ensure Google API key is valid
- Google Gemini AI
- LangChain & LangGraph
- Flask Framework
- AWS
Made with ❤️ for English learners