Vaidya Voice is a revolutionary multilingual, voice-first telemedicine application designed to bridge the healthcare gap in rural India. It enables users to describe their medical symptoms in their native language (Hindi, Bengali, or English) through voice commands, processes the information using advanced AI, and provides medical triage advice through spoken responses.
- Tap to Speak: No typing required - simply tap the microphone and describe your symptoms
- Natural Conversation: The AI understands and responds naturally, making healthcare accessible for everyone
- Smart Recording: Automatic audio capture with visual feedback
- Hindi (ΰ€Ήΰ€Ώΰ€¨ΰ₯ΰ€¦ΰ₯): Complete support for Hindi language with proper Devanagari script responses
- Bengali (বাΰ¦ΰ¦²ΰ¦Ύ): Full Bengali language support with Bengali script responses
- English: Native English language support
- Automatic Translation: The AI instantly translates and understands input in any supported language
- Powered by Meta Llama 3: Uses state-of-the-art AI model via Groq API for medical analysis
- Symptom Analysis: Analyzes symptoms and asks relevant follow-up questions
- Medical Triage: Provides appropriate medical advice and recommendations
- Context Awareness: Remembers previous interactions in the conversation for better understanding
- Live Video Streaming: Real-time WebRTC video streaming for visual examination
- Static Image Upload: Option to upload photos of injuries, rashes, or visible symptoms
- AI Vision Analysis: Advanced computer vision analyzes visual symptoms alongside voice descriptions
- MongoDB Atlas Integration: Secure cloud storage for patient consultation history
- Conversation Context: Maintains context throughout the consultation session
- Data Privacy: All patient data is securely stored and encrypted
- Text-to-Speech: The AI doctor speaks responses back to the patient
- Native Language Voices: Uses appropriate voice synthesis for each language
- Clear Audio: High-quality audio output for better understanding
- Framework: React Native with Expo
- Navigation: Expo Router for seamless navigation
- Audio Processing: Expo AV for recording and playback
- Speech Recognition: Expo Speech for text-to-speech
- Camera Integration: Expo Camera and Image Picker
- Real-time Communication: WebRTC for live video streaming
- UI Components: Beautiful, responsive design with animations
- Framework: Python FastAPI for high-performance API
- AI Processing: Groq API with Llama 3 models
- Audio Processing: SpeechRecognition, Pydub, FFmpeg
- Database: MongoDB Atlas with Beanie ODM
- Real-time Communication: WebRTC for video streaming
- File Handling: Multipart file uploads and processing
- Speech-to-Text: Groq Whisper for accurate transcription
- Medical AI: Meta Llama 3 for medical diagnosis
- Vision AI: Llama 4 Scout for image analysis
- Translation: Deep Translator for language processing
- Node.js: Version 18+ (for React Native development)
- Python: Version 3.10+ (for backend development)
- npm/yarn: Package manager for Node.js dependencies
- Git: For version control
- MongoDB Atlas: Free tier cloud database account
- Groq API: Free tier API key for AI services
- Expo Go App: For mobile testing (Android/iOS)
- FFmpeg: Required for audio format conversion
- Android Studio: For Android development (optional)
- Xcode: For iOS development (macOS only, optional)
git clone https://github.com/your-username/Project-Vaidya-Voice.git
cd Project-Vaidya-Voicecd backend
python -m venv venv
# Activate virtual environment:
# Windows:
venv\Scripts\activate
# Mac/Linux:
source venv/bin/activatepip install fastapi uvicorn groq pymongo beanie speechrecognition gTTS pydub python-multipart python-dotenv deep-translator motor aiofiles aiortc PillowCreate a .env file in the backend/ directory:
# MongoDB Atlas Connection String
MONGO_URL=mongodb+srv://<username>:<password>@cluster0.xxxxx.mongodb.net/?retryWrites=true&w=majority
# Groq API Key (Get from https://console.groq.com/)
GROQ_API_KEY=gsk_your_groq_key_here# Start the backend server
uvicorn app.main:app --reload --host 0.0.0.0 --port 8000cd ../Vaidya-Voice
npm install
# or
yarn installOpen Vaidya-Voice/app/index.tsx and update the backend URL:
// Find this line (around line 127):
const backendUrl = "http://YOUR_LOCAL_IP:8000/api/analyze-voice";
// Replace YOUR_LOCAL_IP with your computer's local IP address
// Find your IP by running:
# Windows: ipconfig
# Mac/Linux: ifconfigAlso update the WebRTC URL in the same file:
// Find this line (around line 201):
const response = await fetch('http://YOUR_LOCAL_IP:8000/offer', {- Ensure your mobile device and development machine are on the same WiFi network
- Install Expo Go app on your mobile device (from App Store/Play Store)
- Windows Firewall should allow Python/Uvicorn connections
# Terminal 1: Navigate to backend directory
cd backend
# Activate virtual environment (if not already active)
# Windows: venv\Scripts\activate
# Mac/Linux: source venv/bin/activate
# Start the FastAPI server
uvicorn app.main:app --reload --host 0.0.0.0 --port 8000Important: The --host 0.0.0.0 flag is crucial for mobile device connectivity.
# Terminal 2: Navigate to frontend directory
cd Vaidya-Voice
# Start the Expo development server
npm start
# or
expo start- Open Expo Go app on your mobile device
- Scan the QR code displayed in the terminal
- The app will automatically load and connect to your backend
- Select Language: Choose between English, Hindi, or Bengali
- Tap Microphone: Press and hold the microphone button
- Describe Symptoms: Speak clearly about your medical concerns
- Release to Process: Release the microphone to send for analysis
- Listen to Response: The AI doctor will provide advice in your language
- Go Live: Tap the red video camera button to start live streaming
- Point Camera: Aim your camera at the affected area (injury, rash, etc.)
- Speak Symptoms: While streaming, describe your symptoms via voice
- Receive Analysis: AI analyzes both video and voice for comprehensive diagnosis
- End Session: Tap the red button again to stop streaming
- Select Camera: Tap the camera icon (not available during live mode)
- Choose Source: Take new photo or select from gallery
- Capture Image: Take or select an image of the affected area
- Voice Description: Record voice description of symptoms
- Get Analysis: AI analyzes both image and voice input
- Check IP Address: Ensure backend URL in
index.tsxmatches your local IP - Network Connection: Verify phone and laptop are on same WiFi
- Firewall: Allow Python/Uvicorn through Windows Defender Firewall
- Backend Status: Ensure backend server is running on port 8000
- Permissions: Grant microphone permissions to Expo Go app
- Audio Settings: Check if phone's microphone is enabled
- App Restart: Restart the Expo Go app
- Permissions: Grant camera permissions to Expo Go app
- WebRTC Issues: Restart both backend and frontend servers
- Network Stability: Ensure stable WiFi connection
- Clear Speech: Speak clearly in the selected language
- Language Selection: Ensure correct language is selected before speaking
- Background Noise: Minimize background noise during recording
- API Keys: Verify Groq API key is valid and has credits
- Internet Connection: Check stable internet connectivity
- Server Logs: Check backend terminal for error messages
# Check backend logs in terminal
# Look for these messages:
# --- βοΈ Connected to MongoDB Cloud Successfully ---
# π Incoming WebRTC Connection...
# π§ Thinking... Patient said: [symptoms]- Expo Developer Tools: Open browser at
http://localhost:19002 - Console Logs: Check browser console for JavaScript errors
- Network Tab: Monitor API requests and responses
Project-Vaidya-Voice/
βββ π± Vaidya-Voice/ # React Native Mobile App
β βββ app/
β β βββ index.tsx # Main application screen
β β βββ _layout.tsx # App layout and navigation
β βββ assets/ # Images, icons, and fonts
β βββ package.json # Node.js dependencies
β βββ app.json # Expo configuration
β βββ README.md # Frontend-specific docs
β
βββ π§ backend/ # Python FastAPI Server
β βββ app/
β β βββ api/
β β β βββ endpoints/
β β β βββ diagnosis.py # Main API endpoints
β β βββ core/
β β β βββ database.py # MongoDB connection
β β β βββ state.py # WebRTC state management
β β β βββ config.py # Configuration settings
β β βββ models/
β β β βββ consultation.py # Database models
β β βββ services/
β β βββ medical_brain.py # AI diagnosis logic
β β βββ audio_transcription.py # Speech-to-text
β β βββ translator.py # Language translation
β βββ requirements.txt # Python dependencies
β βββ .env # Environment variables
β βββ venv/ # Python virtual environment
β
βββ π README.md # This file
- Encrypted Storage: All patient data is encrypted in MongoDB Atlas
- Secure Transmission: HTTPS/WSS protocols for data transmission
- Local Processing: Audio files are temporarily stored and deleted after processing
- No Personal Data: No collection of personally identifiable information
- Environment Variables: Sensitive keys stored in
.envfiles - CORS Configuration: Proper cross-origin resource sharing setup
- Input Validation: Server-side validation for all inputs
- Rate Limiting: Protection against API abuse
- Fork Repository: Create a fork of the project
- Create Branch: Make a feature branch from main
- Make Changes: Implement your changes with proper testing
- Test Thoroughly: Test on both Android and iOS platforms
- Submit PR: Create a pull request with detailed description
- Python: Follow PEP 8 style guidelines
- TypeScript: Use ESLint configuration provided
- Comments: Add meaningful comments for complex logic
- Documentation: Update README for new features
This project is licensed under the MIT License - see the LICENSE file for details.
- GitHub Issues: Report bugs and request features via GitHub Issues
- Documentation: Refer to this README and inline code comments
- Community: Join our Discord community for real-time support
- π©Ί Vital Signs Integration: Connect with wearable health devices
- π Medicine Recognition: AI-powered medication identification
- π¨ββοΈ Doctor Consultation: Connect with real doctors for complex cases
- π Health Analytics: Personalized health tracking and insights
- π More Languages: Support for additional Indian languages
- π§ Offline Mode: Basic functionality without internet
- β‘ Performance: Faster response times and reduced latency
- π¨ UI/UX: Enhanced user interface and experience
- π Enhanced Security: Advanced security features and compliance
- Groq: For providing powerful AI models and API services
- MongoDB Atlas: For reliable cloud database infrastructure
- Expo: For excellent React Native development platform
- Meta AI: For the Llama 3 and Llama 4 models
- Open Source Community: For the amazing tools and libraries
Made with β€οΈ for Rural Healthcare in India
Vaidya Voice - Bringing Quality Healthcare to Every Village