-
Notifications
You must be signed in to change notification settings - Fork 10
Expand file tree
/
Copy pathTODO.org
More file actions
516 lines (435 loc) · 24.2 KB
/
Copy pathTODO.org
File metadata and controls
516 lines (435 loc) · 24.2 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
#+TITLE: Tasks to do for simulflow
#+startup: indent content
* TODO Make new scenario node trigger the activity detection timer as well
* TODO Create Mute Filter processor that makes transport in mute the user based on certain conditions
* DONE Add standard =vad= keywords that are supported by simulflow by default (currently just =silero/vad=)
CLOSED: [2025-08-27 Wed 13:23]
:LOGBOOK:
CLOCK: [2025-08-27 Wed 10:08]--[2025-08-27 Wed 13:23] => 3:15
:END:
* DONE Change audio utils for conversion into specific encoding change (ulaw->pcm, pcm->ulaw etc) and resampling (same encoding, different sample rate))
CLOSED: [2025-09-02 Tue 19:56]
* DONE Standardize serializers/deserializers to have within the pipeline 16kHz PCM audio
CLOSED: [2025-09-02 Tue 19:56]
* TODO [#A] Add TTFT metric
* TODO [#A] Add support for telnyx transport
* TODO [#A] Add LLM usage metrics based on chunks responses [[https://platform.openai.com/docs/api-reference/chat/object#chat/object-usage][API docs for usage]]
* TODO [#A] Implement background noise filtering with [[https://docs.pipecat.ai/guides/features/krisp][krisp.ai]]
* TODO [#A] Add pipeline interruptions :mvp:
Our task for today is to finalize the interruption flow to work correctly. Don't
edit anything, we just want to plan the implementation. The interruption flow
consists of this:
- if the user speaks, aka a user-speech-start frame is emitted, the transport
realtime output is interrupted from sending out realtime out frames
At the same time, transport-out (realtime out process) discards new
raw-audio-out frames until it receives a user-speech-stop frame
- the llm process (openai.clj for example) when it receives a user-speech-start
frame, if there is a in-flight request for generating responses, that request
should be discarded - here we need to plan ahead for request interruption as
that is not yet supported. We use hato which is a clojure wrapper over the
Java 11 HTTP Client
- the text-to-speech process needs to drop its audio buffer if it is
accumulating an audio response, and also maintain a :pipeline/interrupted?
state. New results from the websocket connection will be dropped while
:pipeline/interrupted? is true. This flag is toggled true on
:user-speech-start and ends on :user-speech-stop frames
Actually sorry, it is not on :user-speech-stop/start frames but on
=::control-interrupt-start= =::control-interrupt-stop= that these things occur.
- simulflow uses a [[file:src/simulflow/processors/llm_context_aggregator.clj::(defn- llm-sentence-assembler-impl][sentence assembler]] that assembles streaming llm tokens until
it has a full sentence and sends it forward as a speak frame. Usually this
goes to the TTS process and to the context aggregator. This process also needs
to drop all accumulated buffer for a partial sentence and not assemble anymore
when pipeline is interrupted
- [[file:src/simulflow/processors/llm_context_aggregator.clj::(def context-aggregator][Context aggregator]] takes all speak frames and keeps a full ledger of the
conversation so far. The trouble is that the context aggregator appends a new
LLM sentence without any concern if the user heard it. (audio equivalent of
that sentence was played back to the user). Text to speech processors like
Elevenlabs return also char timing information. Simulflow maintains this info
with [[file:src/simulflow/frame.clj::(def XiAlignment \[:map][vendor specific frames]]. Currently, they are not used in processing. As an
example, they were used by a user of simulflow to track when a specific
sentence was said. We could normalize the char alignment into a standard
format and make audio chunking at least fit specific chars. This way, the
transport out could emit a "heard" frame - a frame that denotes that specific
word was heard by the user so it can be safely appended to the context even if
the full sentence wasn't heard.
Here's another description of this process from the perspective of another
popular realtime ai orchestration platform: pipecat
** How pipecat handles interruptions
In pipecat, transport contains both [[file:~/workspace/pipecat/src/pipecat/transports/base_input.py::class BaseInputTransport(FrameProcessor):][in]] & out. When audio comes in, each chunk is
sent to a VAD analizer like Silero and optionally a Smart turn detection model.
Based on these analyzers the BaseTransportInput keeps a =speaking= flag. If
there is a smart turn detector, it's logic for wether the user is
speaking/stopped speaking will take precedence, otherwise the VAD will be the
source of truth.
When the VAD state changes, (ex: speaking -> not speaking), the pipeline emits
=VADUserStartedSpeakingFrame=/=VADUserStoppedSpeakingFrame=. This happens
regardless in order to have good observation on the pipeline.
The transport input has a role into starting interruptions *if* there isn't a
interruption strategy created. If one or more *interruption* strategies exists,
the triggering of the interruption logic is defered to use UserContextAggregator
that has access to the current transcriptions coming in.
Interruption strategies differ from normal VAD or Turn Taking because they can
implement custom logic like: Interrupt the user only if the user said 3 words.
The context aggregator will send a =BotInterrupt= frame to transport in which
will send a =StartInterruptionFrame=.
*** Interruption flow in pipecat (very similar conceptually for simulflow)
1. When either the BaseTransportInput or the UserContextAggregator deems an
interruption should start, they emit a frame to do so. BaseTransportInput emits
a =StartInterruptionFrame= and UserContextAggregator emits a
=BotInterruptionFrame= which is sent to the BaseTransportInput who upon
receiving this frame, emits a =StartInterruptionFrame=. This frame does the
following:
- Tells the LLM to cancel in-flight requests and current streaming tokens
- Tells TTS processors to cancel in-flight reqeusts and clear their accumulators
- Tells TransportOut to stop sending AudioOut frames and clear the current
playback queue
- (Possibly) tells BotContextAggregator to drop current sentence assembled or
just cut it short as the user only heard a part of it.
2. The pipeline is now in a interrupted state (relevant for TransportOut because
it drops any new AudioOut frames until it receives a =StopInterruptionFrame=)
3. When user is deemed to have stopped speaking (by either VAD or Turn taking
model) a =UserStoppedSpeaking= frame is sent. Which will trigger a
=StopInterruptionFrame= if the pipeline supports interuption. This frame will
either be sent by the TransportBaseInput.
Important mention here: The LLM, TTS don't keep a =speaking= flag in this
period as they don't care about this state. They only drop their current
activity when a =StartInterruptionFrame= is received but don't handle a
=StopInterruptionFrame= at all.
** Differences between pipecat and simulflow
1. (I think) simulflow TTS processors whould keep a =:pipeline/interrupted?=
state because when the processor receives a =speak-frame=, it sends it on the
websocket connection to the actual TTS provider that may send one or more
events back that need to be accumulated to construct the full audio
eequivalent of the text from the =speak-frame=. Therefore we keep the
=pipeline/interrupted?= flag so when new data is received on the websocket
the processor drops them.
2. We need a way to clear the "playback queue". Currently the playback queue is
represented by the [[file:src/simulflow/transport/out.clj::audio-write-ch (a/chan 1024)\]][audio-write-channel]] defined. There is a [[file:src/simulflow/async.clj::(defn drain-channel!][drain-channel!]]
function which will work but we need to introduce two channels to communicate
with the [[file:src/simulflow/transport/out.clj::(vthread-loop \[\]][process running in a vthread]] that sends audio to out. One for
commands to drain audio, and one on which to take audio from (the current
existing one)
3. Pipecat uses a bidirectional queue system between processors:
Transport in <-> Transcriptor <-> Context Aggregator <-> LLM <-> TTS <->
Transport out
Processors send frames either forward or backward (upstream or downstream). They
also have two different queues on each direction, to account for normal frames
and system frames, because you want system frames (like interrupt-start) to be
handled immediately, even if on the queue you have more frames to handle.
This system makes it easy to send system frames and be sure they are received by
all interested processors. The downside is that all processors need to process
all frames, by at a minimum sending them to the next process in the defined
direction.
Simulflow defines processes in a directed graph that can have cycles, so as a
user you need to define exactly the connections needed for processors to
interract correctly.
This is more efficient, in the sense that processors need to handle only the
frames they care about and don't have to worry about sending frames forward
however this complicates system frames propagation. and you end up with large
edge definitions. Example:
#+begin_src clojure
(defn phone-flow
"This example showcases a voice AI agent for the phone. Phone audio is usually
encoded as MULAW at 8kHz frequency (sample rate) and it is mono (1 channel)."
[{:keys [llm-context extra-procs in out extra-conns language vad-analyser]
:or {llm-context {:messages [{:role "system"
:content "You are a helpful assistant "}]}
extra-procs {}
language :en
extra-conns []}}]
(let [chunk-duration-ms 20]
{;; procs are the processes involved in the pipeline. They are nodes in the graph
:procs
(u/deep-merge
{:transport-in {:proc transport-in/twilio-transport-in
:args {:transport/in-ch in
:vad/analyser vad-analyser}}
:transcriptor {:proc asr/deepgram-processor
:args {:transcription/api-key (secret [:deepgram :api-key])
:transcription/interim-results? true
:transcription/punctuate? false
:transcription/vad-events? false
:transcription/smart-format? true
:transcription/model :nova-2
:transcription/utterance-end-ms 1000
:transcription/language language}}
:context-aggregator {:proc context/context-aggregator
:args {:llm/context llm-context
:aggregator/debug? false}}
:llm {:proc llm/openai-llm-process
:args {:openai/api-key (secret [:openai :new-api-sk])
:llm/model :gpt-4.1-mini}}
:assistant-context-assembler {:proc context/assistant-context-assembler
:args {:debug? false}}
:llm-sentence-assembler {:proc context/llm-sentence-assembler}
:tts {:proc tts/elevenlabs-tts-process
:args {:elevenlabs/api-key (secret [:elevenlabs :api-key])
:elevenlabs/model-id "eleven_flash_v2_5"
:elevenlabs/voice-id (secret [:elevenlabs :voice-id])
:voice/stability 0.5
:voice/similarity-boost 0.8
:voice/use-speaker-boost? true
:pipeline/language language
:audio.out/sample-rate 8000
:audio.out/encoding :ulaw}}
:audio-splitter {:proc transport/audio-splitter
:args {:audio.out/duration-ms chunk-duration-ms
:audio.out/sample-size-bits 8}}
:transport-out {:proc transport-out/realtime-out-processor
:args {:audio.out/chan out
:audio.out/sending-interval chunk-duration-ms}}
:activity-monitor {:proc activity-monitor/process
:args {::activity-monitor/timeout-ms 5000}}}
extra-procs)
;; :conns are the edges of the graph. The :out, :in, :sys-out, :sys-in are
;; the channels each processor defines. They are described in the 0 arity
;; version of the processor or the describe function
:conns (concat
[[[:transport-in :out] [:transcriptor :in]]
[[:transcriptor :out] [:context-aggregator :in]]
[[:transport-in :sys-out] [:context-aggregator :sys-in]]
[[:transport-in :sys-out] [:transport-out :sys-in]]
[[:context-aggregator :out] [:llm :in]]
;; Aggregate full context
[[:llm :out] [:assistant-context-assembler :in]]
[[:assistant-context-assembler :out] [:context-aggregator :in]]
;; Assemble sentence by sentence for fast speech
[[:llm :out] [:llm-sentence-assembler :in]]
[[:llm-sentence-assembler :out] [:tts :in]]
[[:tts :out] [:audio-splitter :in]]
[[:audio-splitter :out] [:transport-out :in]]
;; Activity detection
[[:transport-out :sys-out] [:activity-monitor :sys-in]]
[[:transport-in :sys-out] [:activity-monitor :sys-in]]
[[:transcriptor :sys-out] [:activity-monitor :sys-in]]
[[:activity-monitor :out] [:context-aggregator :in]]
[[:activity-monitor :out] [:tts :in]]]
extra-conns)}))
#+end_src
There is a [[file:src/simulflow/processors/system_frame_router.clj::(ns simulflow.processors.system-frame-router][system frame router]] process that can be used but there is a caveat:
the system frame router, fans out all of the system frames he receives to his
=:sys-out= channel but all the processors that define an edge from his output,
need to ensure to not send system frames forward because it will cause an
infinite loop. Just to be mentioned, all processors that emit system frames,
will receive back the same system frame from the system route simply by the
nature of the setup.
** TODO Make assistant context aggregator support interrupt :mvp:
** Adding to context only what the user has heard before interruption happened.
We'll need to keep a =context-id= or =sentence-id= for each resulting
=audio-out-raw= frame from TTS service. The original =sentence-id= will be kept on
each =audio-out-raw= resulting from [[file:src/simulflow/transport.clj::(def audio-splitter][audio splitter]] to provide realtime.
The TTS processor will output a word-timestamp frame with the same =sentence-id=
so it can be matched when playback happens.
The [[file:src/simulflow/transport/out.clj::(def realtime-out-processor][transport-out]] processor will receive the realtime =audio-out-raw= frames and
keep the =sentence-id= in local state until it has been played back:
1. Depending which one comes it first: =word-timestamp= frame or the
=audio-out-raw=, the state will keep a map of ={sentence-id:
{word-timestamps started-playback?}}=
2. When the first audio frame is played back, started-playback? is turned to true
3. A new command will be added: =:command/output-words=. The handler from the
init processor will receive the =word-timestamp= data and wait based on the
computed end time of each word to send back to the transform a =word-played=
msg which will output a =WordPlayedFrame= or a =word-heard= that has the
=sentence-id=, =word= and if it marks the sentence end.
4. The LLMSentenc
* TODO Add support for first message greeting in the pipeline :mvp:
* TODO Add support for [[https://github.com/fixie-ai/ultravox][ultravox]]
* TODO Use [[https://github.com/taoensso/trove][trove]] as a logging facade so we don't force users to use telemere for logging
* TODO Add support for openai realtime API
* TODO Research webrtc support
* TODO research [[https://github.com/phronmophobic/clj-media][clojure-media]] for dedicated ffmpeg support for media conversion
* TODO Make a helper to create easier connections between processors
#+begin_src clojure
(def phone-flow
(simulflow/create-flow {:language :en
:transport {:mode :telephony
:in (input-channel)
:out (output-channel)}
:transcriptor {:proc asr/deepgram-processor
:args {:transcription/api-key (secret [:deepgram :api-key])
:transcription/model :nova-2}}
:llm {:proc llm/openai-llm-process
:args {:openai/api-key (secret [:openai :new-api-sk])
:llm/model "gpt-4o-mini"}}
:tts {:proc tts/elevenlabs-tts-process
:args {:elevenlabs/api-key (secret [:elevenlabs :api-key])
:elevenlabs/model-id "eleven_flash_v2_5"}}}))
#+end_src
* TODO Add Gladia as a transcription provider
Some code from another project
#+begin_src clojure
;;;;;;;;; Gladia ASR ;;;;;;;;;;;;;
;; :frames_format "base64"
;; :word_timestamps true})
(def ^:private gladia-url "wss://api.gladia.io/audio/text/audio-transcription")
;; this may be outdated
(def ^:private asr-configuration {:x_gladia_key api-key
:sample_rate 8000
:encoding "WAV/ULAW"
:language_behaviour "manual"
:language "romanian"})
(defn transcript?
[m]
(= (:event m) "transcript"))
(defn final-transcription?
[m]
(and (transcript? m)
(= (:type m) "final")))
(defn partial-transcription?
[m]
(and (transcript? m)
(= (:type m) "partial")))
(defrecord GladiaASR [ws asr-chan]
ASR
(send-audio-chunk [_ data]
(send! ws {:frames (get-in data [:media :payload])} false))
(close! [_]
(ws/close! ws)))
(defn- make-gladia-asr!
[{:keys [asr-text]}]
;; TODO: Handle reconnect & errors
(let [ws @(websocket gladia-url
{:on-open (fn [ws]
(prn "Open ASR Stream")
(send! ws asr-configuration)
(u/log ::gladia-asr-connected))
:on-message (fn [_ws ^HeapCharBuffer data _last?]
(let [m (json/parse-if-json (str data))]
(u/log ::gladia-msg :m m)
(when (final-transcription? m)
(u/log ::gladia-asr-transcription :sentence (:transcription m) :transcription m)
(go (>! asr-text (:transcription m))))))
:on-error (fn [_ e]
(u/log ::gladia-asr-error :exception e))
:on-close (fn [_ code reason]
(u/log ::gladia-asr-closed :code code :reason reason))})]
(->GladiaASR ws asr-text)))
#+end_src
* TODO Add openai text to speech
#+begin_src clojure
(require '[wkok.openai-clojure.api :as openai])
(defn openai
"Generate speech using openai"
([input]
(openai input {}))
([input config]
(openai/create-speech (merge {:input input
:voice "alloy"
:response_format "wav"
:model "tts-1"}
config)
{:version :http-2 :as :stream})))
(defn tts-stage-openai
[sid in]
(a/go-loop []
(let [sentence (a/<! in)]
(when-not (nil? sentence)
(append-message! sid "assistant" sentence)
(try
(let [sentence-stream (-> (tts/openai sentence) (io/input-stream))
ais (AudioSystem/getAudioInputStream sentence-stream)
twilio-ais (audio/->twilio-phone ais)
buffer (byte-array 256)]
(loop []
(let [bytes-read (.read twilio-ais buffer)]
(when (pos? bytes-read)
(twilio/send-msg! (sessions/ws sid)
sid
(e/encode-base64 buffer))
(recur)))))
(catch Exception e
(u/log ::tts-stage-error :exception e)))
(recur)))))
#+end_src
* TODO Add rime ai text to speech
#+begin_src clojure
(def ^:private rime-tts-url "https://users.rime.ai/v1/rime-tts")
(defn rime
"Generate speech using rime-ai provider"
[sentence]
(-> {:method :post
:url rime-tts-url
:as :stream
:body (json/->json-str {:text sentence
:reduceLatency false
:samplingRate 8000
:speedAlpha 1.0
:modelId "v1"
:speaker "Colby"})
:headers {"Authorization" (str "Bearer " rime-api-key)
"Accept" "audio/x-mulaw"
"Content-Type" "application/json"}}
(client/request)
:body))
(defn rime-async
"Generate speech using rime-ai provider, outputs results on a async
channel"
[sentence]
(let [stream (-> (rime sentence)
(io/input-stream))
c (a/chan 1024)]
(au/input-stream->chan stream c 1024)))
(defn tts-stage
[sid in]
(a/go-loop []
(let [sentence (a/<! in)]
(when-not (nil? sentence)
(append-message! sid "assistant" sentence)
(try
(let [sentence-stream (-> (tts/rime sentence) (io/input-stream))
buffer (byte-array 256)]
(loop []
(let [bytes-read (.read sentence-stream buffer)]
(when (pos? bytes-read)
(twilio/send-msg! (sessions/ws sid)
sid
(e/encode-base64 buffer))
(recur)))))
(catch Exception e
(u/log ::tts-stage-error :exception e)))
(recur)))))
#+end_src
* TODO Add support for [[https://talon.wiki/][Talon]] STT
* DONE Add float32 conversion that is fast to use with VAD or turn detection models
CLOSED: [2025-08-12 Tue 17:57]
* DONE Add support for Silero VAD
CLOSED: [2025-08-12 Tue 17:56] DEADLINE: <2025-01-20 Mon 20:00>
:LOGBOOK:
CLOCK: [2025-01-13 Mon 07:54]--[2025-01-13 Mon 08:19] => 0:25
:END:
* DONE Add support for google gemini
CLOSED: [2025-05-13 Tue 11:29]
* DONE Add local transport (microphone + speaker out)
CLOSED: [2025-05-13 Tue 11:30]
:LOGBOOK:
CLOCK: [2025-02-06 Thu 08:07]--[2025-02-06 Thu 08:32] => 0:25
:END:
* DONE Implement diagram flows into vice-fn
CLOSED: [2025-05-13 Tue 11:30]
:LOGBOOK:
CLOCK: [2025-02-02 Sun 10:39]--[2025-02-02 Sun 11:04] => 0:25
CLOCK: [2025-02-02 Sun 07:31]--[2025-02-02 Sun 07:56] => 0:25
CLOCK: [2025-02-01 Sat 11:10]--[2025-02-01 Sat 11:42] => 0:32
CLOCK: [2025-02-01 Sat 05:26]--[2025-02-01 Sat 05:51] => 0:25
CLOCK: [2025-01-31 Fri 07:12]--[2025-01-31 Fri 07:37] => 0:25
CLOCK: [2025-01-31 Fri 06:32]--[2025-01-31 Fri 06:57] => 0:25
:END:
This means implementing flow diagrams
#+begin_src clojure
{:initial-node :start
:nodes
{:start {:role_messages [{:role :system
:content "You are an order-taking assistant. You must ALWAYS use the available functions to progress the conversation. This is a phone conversation and your responses will be converted to audio. Keep the conversation friendly, casual, and polite. Avoid outputting special characters and emojis."}]
:task_messages [{:role :system
:content "For this step, ask the user if they want pizza or sushi, and wait for them to use a function to choose. Start off by greeting them. Be friendly and casual; you're taking an order for food over the phone."}]}
:functions [{:type :function
:function {:name :choose_sushi
:description "User wants to order sushi. Let's get that order started"
}}]
}}
#+end_src
** DONE Implement pre-actions & post actions
CLOSED: [2025-05-13 Tue 11:30]
:LOGBOOK:
CLOCK: [2025-02-03 Mon 09:35]--[2025-02-03 Mon 10:00] => 0:25
:END: