A Turkish-speaking desktop voice assistant built with Python & Tkinter — real-time conversational AI (Gemini Live) with automatic fallback, plus Spotify, smart bulb, WhatsApp, browser, file and system control, AI image/code generation, and more.
Python ve Tkinter ile geliştirilmiş, Türkçe konuşan bir masaüstü sesli asistanı — gerçek zamanlı yapay zeka sohbeti (Gemini Live, otomatik yedekli), Spotify, akıllı ampul, WhatsApp, tarayıcı, dosya ve sistem kontrolü, yapay zeka ile resim/kod üretimi ve daha fazlası.
🚧 This is version 1.1 — actively under development, more features and improvements are on the way. 🚧 Bu, 1.1 sürümüdür — aktif olarak geliştiriliyor, yeni özellikler ve iyileştirmeler yolda.
Vera is a Turkish-language desktop voice assistant written in Python. It wakes up on a whistle or the word "Vera", holds a real-time, natural, barge-in-capable voice conversation through the Gemini Live API, and shows an animated Tkinter orb while it listens/talks. Beyond conversation, Vera can control your smart bulb, Spotify, the whole system, your browser, files on the desktop, and even send WhatsApp messages — all by voice, with spoken confirmation before anything risky.
- 🔴 Real-time voice engine (Gemini Live) — replaced the old turn-by-turn pipeline with a full-duplex WebSocket session: continuous mic streaming, streamed audio replies, and the ability to interrupt Vera mid-sentence, like a real phone call
- 🛟 Automatic voice fallback chain — if Gemini Live can't be reached after repeated attempts, Vera transparently switches to a Groq-based sequential voice session (Whisper STT + edge-tts) so a quota/network hiccup never leaves the assistant mute
- 🌐 Browser control — open any site or search Google, open/search Netflix, and close a specific tab by name (via Chrome DevTools Protocol when available, with a window-automation fallback)
- 📁 File & folder management — create, update and open files/folders on the Desktop by voice (sandboxed to the Desktop by design, for safety)
- 💻 AI code generation — dictate a script/program and Vera writes it with Claude (Anthropic), automatically saved to a file; falls back to Groq if Claude is unavailable
▶️ Run code by voice — ask Vera to run a file she just wrote (or any file on the Desktop): HTML opens as a live preview in the browser, and Python/JS/batch/PowerShell scripts launch in a visible console window, like hitting F5 in an editor- 🖼️ AI image generation — generates and opens an image from a text prompt via Cloudflare Workers AI (FLUX.1 schnell), falling back to Pollinations.ai
- 📺 YouTube lookup — "what's the latest video from <channel>" — finds it, tells you when it was posted, can open it directly
- 💬 WhatsApp Desktop automation — writes a message to a contact and waits for your spoken "send"/"cancel" before actually sending anything
- 🎯 Smarter wake detection — whistle detection now uses real frequency analysis instead of a volume spike, the wake-word listener captures a short pre-roll buffer so the beginning of "Vera" no longer gets clipped and missed, and it briefly ignores audio right after Vera goes back to sleep so her own goodbye line can't echo back and instantly wake her up again
- ⏳ Safer confirmation flow — shutdown/restart/PowerShell-command confirmations now expire after 60 seconds instead of staying "pending" indefinitely
- 🧠 Persistent memory — Vera remembers people you tell her about (relationship, birthday, notes) in MySQL and can recall them in conversation
- 🗣️ Voice control — wake/sleep and background listening modes, real-time speech recognition and natural Turkish text-to-speech
- 🎵 Spotify control — play/pause, search & play a track, volume, "what's playing", opening the Spotify app, and automatic volume ducking while Vera talks
- 💡 Smart bulb control — Tuya smart bulb on/off, color and brightness, via local network
- 🖥️ System control — shutdown/restart/sleep/lock (with spoken confirmation), volume, screenshots, monitor extend/duplicate/second-only, task manager, notification center, file explorer, and a guarded general PowerShell command runner
- 🚀 App launcher — opens desktop apps by voice command
- ⛅ Weather — current weather for any Turkish city/district via OpenWeatherMap
- 💬 General chat — natural conversation through Gemini Live (or the Groq fallback), with Turkish chat history stored in MySQL
- 🎨 Animated GUI — a minimal, animated Tkinter orb that reacts to listening/speaking state
- Python 3 — core language
- Tkinter — desktop GUI
- Gemini Live API — real-time, full-duplex conversational AI over WebSocket
- Groq API — fallback voice conversation, Whisper speech-to-text, and fallback code generation
- Anthropic Claude API — primary AI code generation
- Cloudflare Workers AI / Pollinations.ai — AI image generation
- YouTube Data API v3 — channel/video lookups
- pygame — audio playback engine
- edge-tts — Turkish text-to-speech (fallback voice mode)
- SpeechRecognition + PyAudio — microphone capture & wake-word speech-to-text
- Spotify Web API — playback control
- TinyTuya — local control of the Tuya smart bulb
- pywin32 + PyAutoGUI — Windows/WhatsApp Desktop/browser UI automation
- MySQL — chat history & people memory
- OpenWeatherMap API — weather data
main.py is the only entry point and stays at the project root; it boots the audio engine and starts the Tkinter GUI. Everything else is split into four packages by responsibility:
gui/— the Tkinter window and animated orb (gui.py)core/— settings/constants (config.py), command routing & tool definitions (commands.py), MySQL access (database.py), text helpers (utils.py)controllers/— everything that controls an external device, app or the OS: Spotify (spotify_controller.py), the smart bulb (bulb_controller.py), Windows (system_control.py), app launching (app_launcher.py), the browser (browser_control.py,web_control.py), the Desktop filesystem (file_control.py), WhatsApp Desktop (whatsapp_controller.py)services/— the real-time Gemini Live voice session (live_voice.py), the Groq fallback voice session (groq_voice.py), wake-word/whistle detection (speech.py), AI code generation (claude_codegen.py), AI image generation (image_gen.py), weather (weather.py), YouTube (youtube.py)
Secrets (API keys, DB password) are read from environment variables via python-dotenv — see Usage below.
Vera/
├── main.py # Entry point
├── gui/ # Tkinter GUI
│ └── gui.py
├── core/ # Settings, command routing, DB, helpers
│ ├── config.py
│ ├── commands.py
│ ├── database.py
│ └── utils.py
├── controllers/ # Spotify, bulb, system, apps, browser, files, WhatsApp
│ ├── spotify_controller.py
│ ├── bulb_controller.py
│ ├── system_control.py
│ ├── app_launcher.py
│ ├── browser_control.py
│ ├── web_control.py
│ ├── file_control.py
│ └── whatsapp_controller.py
├── services/ # Voice engines, code/image gen, weather, YouTube
│ ├── live_voice.py
│ ├── groq_voice.py
│ ├── speech.py
│ ├── claude_codegen.py
│ ├── image_gen.py
│ ├── weather.py
│ └── youtube.py
├── img/ # Screenshots
├── .env.example
├── requirements.txt
├── LICENSE
└── README.md
- Clone the repository
git clone https://github.com/AlperT-Code/VeraAI-Assistant-1.1.git cd VeraAI-Assistant-1.1 - Install dependencies (Python 3.10+ recommended)
pip install -r requirements.txt
- Copy
.env.exampleto.envand fill in the keys for the features you want (all optional — an unset key just disables that feature):cp .env.example .env
- Create a MySQL database matching your
.env(DB_NAME, defaultvera_db). - Run it:
python main.py
- Fork the repository
- Create a new branch
- Commit your changes
- Push the branch
- Open a pull request
This project is licensed under the MIT License.
Vera, Python ile yazılmış, Türkçe konuşan bir masaüstü sesli asistanıdır. Islık ya da "Vera" sözüyle uyanır, Gemini Live API üzerinden gerçek zamanlı, doğal ve araya girilebilir bir sesli sohbet yürütür, dinlerken/konuşurken animasyonlu bir Tkinter küresi gösterir. Sohbetin ötesinde; akıllı ampulünü, Spotify'ını, tüm sistemini, tarayıcını, masaüstündeki dosyaları kontrol edebilir, hatta WhatsApp mesajı gönderebilir — hepsi sesle, riskli olan her şeyden önce sözlü onay alarak.
- 🔴 Gerçek zamanlı ses motoru (Gemini Live) — eski sıra-sıra (tur tur) konuşma akışının yerini tam çift yönlü bir WebSocket oturumu aldı: kesintisiz mikrofon akışı, akışlı ses yanıtları ve Vera'yı cümle ortasında sözünü keserek durdurabilme, gerçek bir telefon görüşmesi gibi
- 🛟 Otomatik sesli sohbet yedeği — Gemini Live art arda denemelere rağmen ulaşılamaz olursa Vera sessizce Groq tabanlı sıralı bir sesli oturuma (Whisper STT + edge-tts) geçer, kota/ağ sorunu asistanı hiç konuşamaz bırakmaz
- 🌐 Tarayıcı kontrolü — bir siteyi açma ya da Google'da arama, Netflix açma/arama, isme göre belirli bir sekmeyi kapatma (mümkünse Chrome DevTools Protokolü, yoksa pencere otomasyonu yedeği)
- 📁 Dosya ve klasör yönetimi — masaüstünde sesle dosya/klasör oluşturma, güncelleme ve açma (güvenlik amacıyla bilinçli olarak Masaüstü ile sınırlı)
- 💻 Yapay zeka ile kod yazımı — bir betik/program isteğini söyle, Vera onu Claude (Anthropic) ile yazıp otomatik olarak bir dosyaya kaydeder; Claude'a ulaşılamazsa Groq'a düşer
▶️ Kodu sesle çalıştırma — Vera'nın az önce yazdığı (ya da masaüstündeki herhangi bir) dosyayı çalıştırmasını iste: HTML dosyaları tarayıcıda canlı önizleme olarak açılır, Python/JS/batch/PowerShell betikleri ise bir editördeki F5'e basmak gibi görünür bir konsol penceresinde çalışır- 🖼️ Yapay zeka ile resim üretimi — bir metin isteğinden Cloudflare Workers AI (FLUX.1 schnell) ile resim üretip açar, olmazsa Pollinations.ai'ye düşer
- 📺 YouTube sorgusu — "<kanal>'ın son videosu ne" — videoyu bulur, ne zaman yüklendiğini söyler, istenirse doğrudan açar
- 💬 WhatsApp Desktop otomasyonu — bir kişiye mesaj yazar, gerçekten göndermeden önce senin sözlü "gönder"/"iptal" onayını bekler
- 🎯 Daha akıllı uyanma algılama — ıslık tespiti artık ses yüksekliği yerine gerçek frekans analizi kullanıyor, uyanma kelimesi dinleyicisi kısa bir ön-bellek tutarak "Vera"nın başının kırpılıp kaçırılmasını önlüyor, ayrıca Vera uykuya döndükten hemen sonra kısa bir süre sesi yok sayarak kendi vedalaşma sözünün yankılanıp onu anında yeniden uyandırmasının önüne geçiyor
- ⏳ Daha güvenli onay akışı — kapatma/yeniden başlatma/PowerShell komutu onayları artık süresiz beklemek yerine 60 saniye sonra kendiliğinden düşüyor
- 🧠 Kalıcı hafıza — Vera bahsettiğin kişileri (yakınlık, doğum günü, notlar) MySQL'de hatırlıyor ve sohbette hatırlayabiliyor
- 🗣️ Sesli kontrol — uyanma/uyuma ve arka planda dinleme modları, gerçek zamanlı konuşma tanıma ve doğal Türkçe seslendirme
- 🎵 Spotify kontrolü — çal/duraklat, şarkı arama & çalma, ses seviyesi, "ne çalıyor", Spotify uygulamasını açma ve Vera konuşurken otomatik ses kısma
- 💡 Akıllı ampul kontrolü — Tuya akıllı ampulü yerel ağ üzerinden açma/kapama, renk ve parlaklık ayarı
- 🖥️ Sistem kontrolü — kapat/yeniden başlat/uyut/kilitle (sözlü onayla), ses seviyesi, ekran görüntüsü, monitör genişlet/kopyala/sadece ikinci ekran, görev yöneticisi, bildirim merkezi, dosya gezgini ve korumalı, genel amaçlı bir PowerShell komut çalıştırıcı
- 🚀 Uygulama açıcı — sesli komutla masaüstü uygulamalarını açar
- ⛅ Hava durumu — OpenWeatherMap üzerinden herhangi bir Türkiye il/ilçesi için güncel hava durumu
- 💬 Genel sohbet — Gemini Live (ya da Groq yedeği) üzerinden doğal sohbet, Türkçe sohbet geçmişi MySQL'de tutulur
- 🎨 Animasyonlu arayüz — dinleme/konuşma durumuna tepki veren minimal, animasyonlu bir Tkinter küresi
- Python 3 — ana dil
- Tkinter — masaüstü arayüzü
- Gemini Live API — WebSocket üzerinden gerçek zamanlı, çift yönlü yapay zeka sohbeti
- Groq API — yedek sesli sohbet, Whisper ile sesten metne dönüşüm ve yedek kod yazımı
- Anthropic Claude API — birincil yapay zeka kod üretimi
- Cloudflare Workers AI / Pollinations.ai — yapay zeka ile resim üretimi
- YouTube Data API v3 — kanal/video sorguları
- pygame — ses çalma motoru
- edge-tts — Türkçe metinden sese dönüşüm (yedek ses modu)
- SpeechRecognition + PyAudio — mikrofon yakalama & uyanma kelimesi sesten metne dönüşümü
- Spotify Web API — oynatma kontrolü
- TinyTuya — Tuya akıllı ampulün yerel kontrolü
- pywin32 + PyAutoGUI — Windows/WhatsApp Desktop/tarayıcı arayüz otomasyonu
- MySQL — sohbet geçmişi & kişi hafızası
- OpenWeatherMap API — hava durumu verisi
main.py, tek giriş noktasıdır ve projenin kök dizininde kalır; ses motorunu başlatır ve Tkinter arayüzünü açar. Geri kalan her şey sorumluluğa göre dört pakete ayrılmıştır:
gui/— Tkinter penceresi ve animasyonlu küre (gui.py)core/— ayarlar/sabitler (config.py), komut yönlendirme & araç tanımları (commands.py), MySQL erişimi (database.py), metin yardımcıları (utils.py)controllers/— harici bir cihazı, uygulamayı veya işletim sistemini kontrol eden her şey: Spotify (spotify_controller.py), akıllı ampul (bulb_controller.py), Windows (system_control.py), uygulama açma (app_launcher.py), tarayıcı (browser_control.py,web_control.py), masaüstü dosya sistemi (file_control.py), WhatsApp Desktop (whatsapp_controller.py)services/— gerçek zamanlı Gemini Live sesli oturumu (live_voice.py), Groq yedek sesli oturumu (groq_voice.py), uyanma kelimesi/ıslık tespiti (speech.py), yapay zeka kod üretimi (claude_codegen.py), yapay zeka resim üretimi (image_gen.py), hava durumu (weather.py), YouTube (youtube.py)
Gizli bilgiler (API anahtarları, DB şifresi) python-dotenv ile ortam değişkenlerinden okunur — aşağıdaki Kullanım bölümüne bakın.
Vera/
├── main.py # Giriş noktası
├── gui/ # Tkinter arayüzü
│ └── gui.py
├── core/ # Ayarlar, komut yönlendirme, DB, yardımcılar
│ ├── config.py
│ ├── commands.py
│ ├── database.py
│ └── utils.py
├── controllers/ # Spotify, ampul, sistem, uygulamalar, tarayıcı, dosya, WhatsApp
│ ├── spotify_controller.py
│ ├── bulb_controller.py
│ ├── system_control.py
│ ├── app_launcher.py
│ ├── browser_control.py
│ ├── web_control.py
│ ├── file_control.py
│ └── whatsapp_controller.py
├── services/ # Ses motorları, kod/resim üretimi, hava durumu, YouTube
│ ├── live_voice.py
│ ├── groq_voice.py
│ ├── speech.py
│ ├── claude_codegen.py
│ ├── image_gen.py
│ ├── weather.py
│ └── youtube.py
├── img/ # Ekran görüntüleri
├── .env.example
├── requirements.txt
├── LICENSE
└── README.md
- Depoyu klonlayın
git clone https://github.com/AlperT-Code/VeraAI-Assistant-1.1.git cd VeraAI-Assistant-1.1 - Bağımlılıkları kurun (Python 3.10+ önerilir)
pip install -r requirements.txt
.env.exampledosyasını.envolarak kopyalayıp kullanmak istediğin özelliklerin anahtarlarını girin (hepsi opsiyoneldir — girilmeyen bir anahtar sadece o özelliği devre dışı bırakır):cp .env.example .env
.envdosyanızla eşleşen bir MySQL veritabanı oluşturun (DB_NAME, varsayılanvera_db).- Çalıştırın:
python main.py
- Projeyi fork'layın
- Yeni bir branch oluşturun
- Değişikliklerinizi commit edin
- Branch'i push edin
- Pull request oluşturun
Bu proje MIT Lisansı ile lisanslanmıştır.

