A smart voice assistant written in Python that integrates Google's Gemini API for conversational intelligence. This assistant can perform web automation tasks, tell jokes, and answer general queries using Generative AI.
- Generative AI Conversations: Uses Google's
gemini-2.5-flash-litemodel to answer complex questions and chat. - YouTube Integration: verbal command "YouTube [video name]" automatically searches and plays videos.
- Wikipedia Search: verbal command "Wikipedia [topic]" reads a summary of the topic.
- Entertainment: Can tell jokes using the
pyjokeslibrary. - Natural Text-to-Speech: Uses
gTTS(Google Text-to-Speech) for high-quality voice output. - Personalized Greeting: Greets the user based on the time of day.
- Python 3.x
- Google Generative AI SDK (Gemini)
- SpeechRecognition (Google Speech API)
- gTTS (Google Text-to-Speech)
- PyWhatKit (YouTube automation)
- Wikipedia API
Note: This script currently uses afplay, which is a command-line audio player native to macOS.
- If you are on Windows or Linux, you will need to modify the
speak()function to useplaysoundoros.system("start ...").
You will need a valid API Key from Google AI Studio.
- Go to [Google AI Studio] (https://aistudio.google.com/).
- Create an API key.
-
Clone the repository:
git clone [https://github.com/yourusername/your-repo-name.git](https://github.com/yourusername/your-repo-name.git) cd your-repo-name -
Install the required dependencies:
pip install speechrecognition pyttsx3 wikipedia pyjokes google-generativeai gTTS pywhatkit pyaudio
(Note:
pyaudiois required for microphone input. On some systems, you may need to installportaudiovia Homebrew or apt-get first). -
Configure your API Key: Open the Python script and replace the placeholder with your actual key:
api_key = "YOUR_GOOGLE_GEN_AI_KEY_HERE"
Run the main script:
python main.py