A task manager with an assistant called Kai that knows what is on your list. Chat or talk to Kai, and it answers from your real tasks and categories.
It started as a take-home for a YUPIX interview: build a todo app with an AI assistant. Most "AI todo" demos bolt a generic chatbot next to the list. The bot has no clue what you have due this week. Kai reads your current tasks and categories before it answers, so "what should I do first?" gets a reply grounded in your actual list. I liked the result enough to keep building it past the interview.
- Create, edit, complete, and delete tasks, with due dates and custom categories
- Drag to reorder tasks, with the order saved per list (pending and completed keep separate positions)
- A monthly calendar view of everything with a due date
- Chat with Kai by text, or talk to it and get a spoken reply, streamed live
- Email verification on signup, so accounts are tied to a real inbox
- JWT auth guarding every task, category, and AI route
The frontend is a Next.js app. The backend is a plain Express server that owns the database, the OpenAI calls, and a WebSocket for the live assistant.
flowchart LR
B[Next.js client] -->|REST /api| E[Express API]
B -.->|WebSocket /ws/realtime| E
E --> M[(MongoDB)]
E -->|chat + realtime| O[OpenAI]
E -->|verification mail| R[Resend]
Kai's live mode runs over a WebSocket at /ws/realtime. A browser cannot attach an Authorization header to a WebSocket. So the first frame the client sends is an auth message carrying the JWT. The server verifies it and loads that user's tasks and categories from MongoDB. Those go straight into the OpenAI Realtime session instructions before any audio flows. From there the server relays incremental events back down the same socket: response.audio_transcript.delta for streamed text, response.audio.delta for PCM16 audio, and Whisper transcripts of what you said. By the time you speak, the model already has your real list in context.
flowchart LR
A["Client opens /ws/realtime"] --> B[Auth frame with JWT]
B --> V[Server verifies JWT]
V --> L[Load tasks + categories from MongoDB]
L --> I[Inject into Realtime session instructions]
I --> S[Stream text + audio deltas back]
Each task carries two numbers, pendingOrder and completedOrder, so it holds a separate spot in the active list and the done list. dnd-kit handles the drag on the client. The resulting order is sent once to PUT /api/tasks/reorder. The server picks the field to update from the task's completion state and writes every new index in one Task.bulkWrite. One round trip, one write, and the order survives a refresh.
Login signs a JWT and returns it in the response body. The client keeps it and sends it as a Bearer token on every API call. An Express middleware verifies the token and attaches the user (minus the password hash). On the Next.js side, a middleware guards /todos by checking for the token, so an unauthenticated visitor lands on login before the page renders.
Frontend: Next.js 15, React 19, TypeScript, Tailwind CSS, shadcn/ui, TanStack Query, Zustand, dnd-kit, react-hook-form with Zod
Backend: Node, Express, MongoDB with Mongoose, ws for the WebSocket
AI: OpenAI (gpt-4-turbo for chat, gpt-4o-realtime-preview for the streamed voice/text assistant)
Email: Resend
Auth: JWT with bcrypt password hashing
kai-tasks/
frontend/ Next.js app: todos, calendar, categories, Kai chat and voice
backend/ Express API, Mongoose models, and the realtime WebSocket service
You need Node 18+, a MongoDB instance, and an OpenAI key. A Resend key is only needed if you want verification emails to actually send.
# backend
cd backend
npm install
cp .env.example .env # fill in the values below
npm run dev # port 5000, plus the WebSocket on the same port
# frontend
cd frontend
npm install
npm run dev # port 3000Backend .env:
MONGODB_URI=mongodb://localhost:27017/todo-app
JWT_SECRET=your-secret-key-here
PORT=5000
OPENAI_API_KEY=sk-your-openai-key
RESEND_API_KEY=re_xxxxxxxxxxxx
EMAIL_FROM=noreply@yourdomain.com
FRONTEND_URL=http://localhost:3000
For a hosted database, point MONGODB_URI at your own MongoDB Atlas cluster. Keep the real string out of the repo.
All REST routes are under http://localhost:5000/api.
Auth: POST /auth/signup, POST /auth/login, POST /auth/verify-email, POST /auth/resend-verification
Tasks: GET /tasks, POST /tasks, PUT /tasks/:id, DELETE /tasks/:id, PUT /tasks/reorder
Categories: GET /categories, POST /categories, PUT /categories/:id, DELETE /categories/:id
AI: POST /ai/chat, and the WebSocket at ws://localhost:5000/ws/realtime
Working. The REST chat and the realtime voice/text assistant both run against live OpenAI models. Voice command parsing that maps speech to app actions (create a task by voice, for example) is still stubbed in the REST controller. Today the spoken interaction is conversational; the command layer is not built yet.