Skip to content

Latest commit

Β 

History

14 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸ₯ Vaidya Voice: AI Health Assistant for Rural India

Vaidya Voice is a revolutionary multilingual, voice-first telemedicine application designed to bridge the healthcare gap in rural India. It enables users to describe their medical symptoms in their native language (Hindi, Bengali, or English) through voice commands, processes the information using advanced AI, and provides medical triage advice through spoken responses.

🌟 Key Features

πŸ—£οΈ Voice-First Interface

  • Tap to Speak: No typing required - simply tap the microphone and describe your symptoms
  • Natural Conversation: The AI understands and responds naturally, making healthcare accessible for everyone
  • Smart Recording: Automatic audio capture with visual feedback

🌏 Multilingual Support

  • Hindi (ΰ€Ήΰ€Ώΰ€¨ΰ₯ΰ€¦ΰ₯€): Complete support for Hindi language with proper Devanagari script responses
  • Bengali (বাংলা): Full Bengali language support with Bengali script responses
  • English: Native English language support
  • Automatic Translation: The AI instantly translates and understands input in any supported language

🧠 Intelligent Medical AI

  • Powered by Meta Llama 3: Uses state-of-the-art AI model via Groq API for medical analysis
  • Symptom Analysis: Analyzes symptoms and asks relevant follow-up questions
  • Medical Triage: Provides appropriate medical advice and recommendations
  • Context Awareness: Remembers previous interactions in the conversation for better understanding

πŸ“Έ Visual Diagnosis Support

  • Live Video Streaming: Real-time WebRTC video streaming for visual examination
  • Static Image Upload: Option to upload photos of injuries, rashes, or visible symptoms
  • AI Vision Analysis: Advanced computer vision analyzes visual symptoms alongside voice descriptions

πŸ’Ύ Cloud-Based Memory

  • MongoDB Atlas Integration: Secure cloud storage for patient consultation history
  • Conversation Context: Maintains context throughout the consultation session
  • Data Privacy: All patient data is securely stored and encrypted

πŸ”Š Audio Response System

  • Text-to-Speech: The AI doctor speaks responses back to the patient
  • Native Language Voices: Uses appropriate voice synthesis for each language
  • Clear Audio: High-quality audio output for better understanding

πŸ› οΈ Technical Architecture

Frontend (Mobile App)

  • Framework: React Native with Expo
  • Navigation: Expo Router for seamless navigation
  • Audio Processing: Expo AV for recording and playback
  • Speech Recognition: Expo Speech for text-to-speech
  • Camera Integration: Expo Camera and Image Picker
  • Real-time Communication: WebRTC for live video streaming
  • UI Components: Beautiful, responsive design with animations

Backend (API Server)

  • Framework: Python FastAPI for high-performance API
  • AI Processing: Groq API with Llama 3 models
  • Audio Processing: SpeechRecognition, Pydub, FFmpeg
  • Database: MongoDB Atlas with Beanie ODM
  • Real-time Communication: WebRTC for video streaming
  • File Handling: Multipart file uploads and processing

AI & Machine Learning

  • Speech-to-Text: Groq Whisper for accurate transcription
  • Medical AI: Meta Llama 3 for medical diagnosis
  • Vision AI: Llama 4 Scout for image analysis
  • Translation: Deep Translator for language processing

πŸ“‹ System Requirements

For Development

  • Node.js: Version 18+ (for React Native development)
  • Python: Version 3.10+ (for backend development)
  • npm/yarn: Package manager for Node.js dependencies
  • Git: For version control

External Services

  • MongoDB Atlas: Free tier cloud database account
  • Groq API: Free tier API key for AI services
  • Expo Go App: For mobile testing (Android/iOS)

System Software

  • FFmpeg: Required for audio format conversion
  • Android Studio: For Android development (optional)
  • Xcode: For iOS development (macOS only, optional)

πŸš€ Installation & Setup Guide

Step 1: Clone the Repository

git clone https://github.com/your-username/Project-Vaidya-Voice.git
cd Project-Vaidya-Voice

Step 2: Backend Setup (The AI Brain)

2.1 Create Virtual Environment

cd backend
python -m venv venv

# Activate virtual environment:
# Windows:
venv\Scripts\activate
# Mac/Linux:
source venv/bin/activate

2.2 Install Python Dependencies

pip install fastapi uvicorn groq pymongo beanie speechrecognition gTTS pydub python-multipart python-dotenv deep-translator motor aiofiles aiortc Pillow

2.3 Configure Environment Variables

Create a .env file in the backend/ directory:

# MongoDB Atlas Connection String
MONGO_URL=mongodb+srv://<username>:<password>@cluster0.xxxxx.mongodb.net/?retryWrites=true&w=majority

# Groq API Key (Get from https://console.groq.com/)
GROQ_API_KEY=gsk_your_groq_key_here

2.4 Test Backend Installation

# Start the backend server
uvicorn app.main:app --reload --host 0.0.0.0 --port 8000

Step 3: Frontend Setup (The Mobile App)

3.1 Install Node.js Dependencies

cd ../Vaidya-Voice
npm install
# or
yarn install

3.2 Update Backend URL

Open Vaidya-Voice/app/index.tsx and update the backend URL:

// Find this line (around line 127):
const backendUrl = "http://YOUR_LOCAL_IP:8000/api/analyze-voice";

// Replace YOUR_LOCAL_IP with your computer's local IP address
// Find your IP by running:
# Windows: ipconfig
# Mac/Linux: ifconfig

3.3 Update WebRTC URL

Also update the WebRTC URL in the same file:

// Find this line (around line 201):
const response = await fetch('http://YOUR_LOCAL_IP:8000/offer', {

πŸƒβ€β™‚οΈ Running the Application

Prerequisites

  • Ensure your mobile device and development machine are on the same WiFi network
  • Install Expo Go app on your mobile device (from App Store/Play Store)
  • Windows Firewall should allow Python/Uvicorn connections

Start the Backend Server

# Terminal 1: Navigate to backend directory
cd backend

# Activate virtual environment (if not already active)
# Windows: venv\Scripts\activate
# Mac/Linux: source venv/bin/activate

# Start the FastAPI server
uvicorn app.main:app --reload --host 0.0.0.0 --port 8000

Important: The --host 0.0.0.0 flag is crucial for mobile device connectivity.

Start the Frontend

# Terminal 2: Navigate to frontend directory
cd Vaidya-Voice

# Start the Expo development server
npm start
# or
expo start

Connect Mobile Device

  1. Open Expo Go app on your mobile device
  2. Scan the QR code displayed in the terminal
  3. The app will automatically load and connect to your backend

πŸ“± How to Use Vaidya Voice

Basic Voice Consultation

  1. Select Language: Choose between English, Hindi, or Bengali
  2. Tap Microphone: Press and hold the microphone button
  3. Describe Symptoms: Speak clearly about your medical concerns
  4. Release to Process: Release the microphone to send for analysis
  5. Listen to Response: The AI doctor will provide advice in your language

Live Video Consultation

  1. Go Live: Tap the red video camera button to start live streaming
  2. Point Camera: Aim your camera at the affected area (injury, rash, etc.)
  3. Speak Symptoms: While streaming, describe your symptoms via voice
  4. Receive Analysis: AI analyzes both video and voice for comprehensive diagnosis
  5. End Session: Tap the red button again to stop streaming

Static Image Upload

  1. Select Camera: Tap the camera icon (not available during live mode)
  2. Choose Source: Take new photo or select from gallery
  3. Capture Image: Take or select an image of the affected area
  4. Voice Description: Record voice description of symptoms
  5. Get Analysis: AI analyzes both image and voice input

πŸ”§ Troubleshooting Guide

Common Issues & Solutions

🚨 "Could not connect to Dr. Vaidya" Error

  • Check IP Address: Ensure backend URL in index.tsx matches your local IP
  • Network Connection: Verify phone and laptop are on same WiFi
  • Firewall: Allow Python/Uvicorn through Windows Defender Firewall
  • Backend Status: Ensure backend server is running on port 8000

🎀 Microphone Not Working

  • Permissions: Grant microphone permissions to Expo Go app
  • Audio Settings: Check if phone's microphone is enabled
  • App Restart: Restart the Expo Go app

πŸ“Ή Camera Not Working

  • Permissions: Grant camera permissions to Expo Go app
  • WebRTC Issues: Restart both backend and frontend servers
  • Network Stability: Ensure stable WiFi connection

🌐 Language Detection Issues

  • Clear Speech: Speak clearly in the selected language
  • Language Selection: Ensure correct language is selected before speaking
  • Background Noise: Minimize background noise during recording

🧠 AI Not Responding

  • API Keys: Verify Groq API key is valid and has credits
  • Internet Connection: Check stable internet connectivity
  • Server Logs: Check backend terminal for error messages

Development Mode Debugging

Backend Debugging

# Check backend logs in terminal
# Look for these messages:
# --- ☁️ Connected to MongoDB Cloud Successfully ---
# πŸ“ž Incoming WebRTC Connection...
# 🧠 Thinking... Patient said: [symptoms]

Frontend Debugging

  • Expo Developer Tools: Open browser at http://localhost:19002
  • Console Logs: Check browser console for JavaScript errors
  • Network Tab: Monitor API requests and responses

πŸ—οΈ Project Structure

Project-Vaidya-Voice/
β”œβ”€β”€ πŸ“± Vaidya-Voice/                 # React Native Mobile App
β”‚   β”œβ”€β”€ app/
β”‚   β”‚   β”œβ”€β”€ index.tsx               # Main application screen
β”‚   β”‚   └── _layout.tsx             # App layout and navigation
β”‚   β”œβ”€β”€ assets/                     # Images, icons, and fonts
β”‚   β”œβ”€β”€ package.json                # Node.js dependencies
β”‚   β”œβ”€β”€ app.json                    # Expo configuration
β”‚   └── README.md                    # Frontend-specific docs
β”‚
β”œβ”€β”€ 🧠 backend/                      # Python FastAPI Server
β”‚   β”œβ”€β”€ app/
β”‚   β”‚   β”œβ”€β”€ api/
β”‚   β”‚   β”‚   └── endpoints/
β”‚   β”‚   β”‚       └── diagnosis.py    # Main API endpoints
β”‚   β”‚   β”œβ”€β”€ core/
β”‚   β”‚   β”‚   β”œβ”€β”€ database.py         # MongoDB connection
β”‚   β”‚   β”‚   β”œβ”€β”€ state.py            # WebRTC state management
β”‚   β”‚   β”‚   └── config.py           # Configuration settings
β”‚   β”‚   β”œβ”€β”€ models/
β”‚   β”‚   β”‚   └── consultation.py     # Database models
β”‚   β”‚   └── services/
β”‚   β”‚       β”œβ”€β”€ medical_brain.py    # AI diagnosis logic
β”‚   β”‚       β”œβ”€β”€ audio_transcription.py # Speech-to-text
β”‚   β”‚       └── translator.py       # Language translation
β”‚   β”œβ”€β”€ requirements.txt            # Python dependencies
β”‚   β”œβ”€β”€ .env                        # Environment variables
β”‚   └── venv/                       # Python virtual environment
β”‚
└── πŸ“š README.md                    # This file

πŸ”’ Security & Privacy

Data Protection

  • Encrypted Storage: All patient data is encrypted in MongoDB Atlas
  • Secure Transmission: HTTPS/WSS protocols for data transmission
  • Local Processing: Audio files are temporarily stored and deleted after processing
  • No Personal Data: No collection of personally identifiable information

API Security

  • Environment Variables: Sensitive keys stored in .env files
  • CORS Configuration: Proper cross-origin resource sharing setup
  • Input Validation: Server-side validation for all inputs
  • Rate Limiting: Protection against API abuse

🀝 Contributing Guidelines

Development Workflow

  1. Fork Repository: Create a fork of the project
  2. Create Branch: Make a feature branch from main
  3. Make Changes: Implement your changes with proper testing
  4. Test Thoroughly: Test on both Android and iOS platforms
  5. Submit PR: Create a pull request with detailed description

Code Standards

  • Python: Follow PEP 8 style guidelines
  • TypeScript: Use ESLint configuration provided
  • Comments: Add meaningful comments for complex logic
  • Documentation: Update README for new features

πŸ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.


πŸ†˜ Support & Contact

Getting Help

  • GitHub Issues: Report bugs and request features via GitHub Issues
  • Documentation: Refer to this README and inline code comments
  • Community: Join our Discord community for real-time support

Emergency Medical Disclaimer

⚠️ IMPORTANT: Vaidya Voice is an AI assistant and NOT a substitute for professional medical care. In case of medical emergencies, please contact emergency services or visit a healthcare facility immediately.


🎯 Future Roadmap

Upcoming Features

  • 🩺 Vital Signs Integration: Connect with wearable health devices
  • πŸ’Š Medicine Recognition: AI-powered medication identification
  • πŸ‘¨β€βš•οΈ Doctor Consultation: Connect with real doctors for complex cases
  • πŸ“Š Health Analytics: Personalized health tracking and insights
  • 🌍 More Languages: Support for additional Indian languages

Technical Improvements

  • πŸ”§ Offline Mode: Basic functionality without internet
  • ⚑ Performance: Faster response times and reduced latency
  • 🎨 UI/UX: Enhanced user interface and experience
  • πŸ” Enhanced Security: Advanced security features and compliance

πŸ™ Acknowledgments

  • Groq: For providing powerful AI models and API services
  • MongoDB Atlas: For reliable cloud database infrastructure
  • Expo: For excellent React Native development platform
  • Meta AI: For the Llama 3 and Llama 4 models
  • Open Source Community: For the amazing tools and libraries

Made with ❀️ for Rural Healthcare in India

Vaidya Voice - Bringing Quality Healthcare to Every Village

About

Vaidya Voice A voice-activated telemedicine assistant designed for rural India. This full-stack application enables users to describe symptoms in their native language (Hindi/English) and receive instant, audible medical advice. Built with React Native (Expo) for the mobile interface, FastAPI for the backend, and powered by Meta's Llama 3 AI for

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages