Skip to content
aenorissPublic

About

A to-do app whose assistant answers from your real list. A Next.js and Express stack feeds your tasks to the OpenAI Realtime API, in text or voice.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

21 Commits

Folders and files

Repository files navigation

Kai Tasks

A task manager with an assistant called Kai that knows what is on your list. Chat or talk to Kai, and it answers from your real tasks and categories.

Why I built it

It started as a take-home for a YUPIX interview: build a todo app with an AI assistant. Most "AI todo" demos bolt a generic chatbot next to the list. The bot has no clue what you have due this week. Kai reads your current tasks and categories before it answers, so "what should I do first?" gets a reply grounded in your actual list. I liked the result enough to keep building it past the interview.

What it does

  • Create, edit, complete, and delete tasks, with due dates and custom categories
  • Drag to reorder tasks, with the order saved per list (pending and completed keep separate positions)
  • A monthly calendar view of everything with a due date
  • Chat with Kai by text, or talk to it and get a spoken reply, streamed live
  • Email verification on signup, so accounts are tied to a real inbox
  • JWT auth guarding every task, category, and AI route

How it works

The frontend is a Next.js app. The backend is a plain Express server that owns the database, the OpenAI calls, and a WebSocket for the live assistant.

flowchart LR
  B[Next.js client] -->|REST /api| E[Express API]
  B -.->|WebSocket /ws/realtime| E
  E --> M[(MongoDB)]
  E -->|chat + realtime| O[OpenAI]
  E -->|verification mail| R[Resend]
Loading

The realtime assistant

Kai's live mode runs over a WebSocket at /ws/realtime. A browser cannot attach an Authorization header to a WebSocket. So the first frame the client sends is an auth message carrying the JWT. The server verifies it and loads that user's tasks and categories from MongoDB. Those go straight into the OpenAI Realtime session instructions before any audio flows. From there the server relays incremental events back down the same socket: response.audio_transcript.delta for streamed text, response.audio.delta for PCM16 audio, and Whisper transcripts of what you said. By the time you speak, the model already has your real list in context.

flowchart LR
  A["Client opens /ws/realtime"] --> B[Auth frame with JWT]
  B --> V[Server verifies JWT]
  V --> L[Load tasks + categories from MongoDB]
  L --> I[Inject into Realtime session instructions]
  I --> S[Stream text + audio deltas back]
Loading

Reordering that survives a refresh

Each task carries two numbers, pendingOrder and completedOrder, so it holds a separate spot in the active list and the done list. dnd-kit handles the drag on the client. The resulting order is sent once to PUT /api/tasks/reorder. The server picks the field to update from the task's completion state and writes every new index in one Task.bulkWrite. One round trip, one write, and the order survives a refresh.

Auth

Login signs a JWT and returns it in the response body. The client keeps it and sends it as a Bearer token on every API call. An Express middleware verifies the token and attaches the user (minus the password hash). On the Next.js side, a middleware guards /todos by checking for the token, so an unauthenticated visitor lands on login before the page renders.

Tech stack

Frontend: Next.js 15, React 19, TypeScript, Tailwind CSS, shadcn/ui, TanStack Query, Zustand, dnd-kit, react-hook-form with Zod Backend: Node, Express, MongoDB with Mongoose, ws for the WebSocket AI: OpenAI (gpt-4-turbo for chat, gpt-4o-realtime-preview for the streamed voice/text assistant) Email: Resend Auth: JWT with bcrypt password hashing

Layout

kai-tasks/
  frontend/    Next.js app: todos, calendar, categories, Kai chat and voice
  backend/     Express API, Mongoose models, and the realtime WebSocket service

Running it

You need Node 18+, a MongoDB instance, and an OpenAI key. A Resend key is only needed if you want verification emails to actually send.

# backend
cd backend
npm install
cp .env.example .env     # fill in the values below
npm run dev              # port 5000, plus the WebSocket on the same port

# frontend
cd frontend
npm install
npm run dev              # port 3000

Backend .env:

MONGODB_URI=mongodb://localhost:27017/todo-app
JWT_SECRET=your-secret-key-here
PORT=5000
OPENAI_API_KEY=sk-your-openai-key
RESEND_API_KEY=re_xxxxxxxxxxxx
EMAIL_FROM=noreply@yourdomain.com
FRONTEND_URL=http://localhost:3000

For a hosted database, point MONGODB_URI at your own MongoDB Atlas cluster. Keep the real string out of the repo.

API

All REST routes are under http://localhost:5000/api.

Auth: POST /auth/signup, POST /auth/login, POST /auth/verify-email, POST /auth/resend-verification Tasks: GET /tasks, POST /tasks, PUT /tasks/:id, DELETE /tasks/:id, PUT /tasks/reorder Categories: GET /categories, POST /categories, PUT /categories/:id, DELETE /categories/:id AI: POST /ai/chat, and the WebSocket at ws://localhost:5000/ws/realtime

Status

Working. The REST chat and the realtime voice/text assistant both run against live OpenAI models. Voice command parsing that maps speech to app actions (create a task by voice, for example) is still stubbed in the REST controller. Today the spoken interaction is conversational; the command layer is not built yet.

About

A to-do app whose assistant answers from your real list. A Next.js and Express stack feeds your tasks to the OpenAI Realtime API, in text or voice.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages