Skip to content

198: Track client-closed requests as 499 in usage logging - #199

Open
JaeYeonLee0621 wants to merge 4 commits into
mainfrom
198-non-responding-requests
Open

JaeYeonLee0621 wants to merge 4 commits into
mainfrom
198-non-responding-requests

Conversation

@JaeYeonLee0621

@JaeYeonLee0621 JaeYeonLee0621 commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor
  • Problem : Users reported slow AI responses and some errors, but the dashboard showed no spike in failed requests.
  • Assumption : When a client cancels the connection, status_code is left as NULL, so these requests don't count as failures on the dashboard.
  • Solution : Save client-cancelled requests with status code 499. The dashboard counts status codes ≥ 400 as failures, so these requests will show up there.

Notes

  • This change needs to be tested in the staging environment first :)
  • This is only one hypothesis. Not every client disconnect is a real error, but during an incident, a rise in client cancellations can be a useful signal that the service is unstable.

ex)

Date Total Request Count 200 500 502 503 504 499 Aborted (NULL)
2026-09-28 25,004 24,309 18 0 0 0 0 465
2026-09-27 32,456 31,612 42 0 0 0 0 120
2026-09-26 33,970 32,852 36 84 1 0 0 82
2026-09-25 38,866 36,253 26 0 0 0 0 1,641
2026-09-24 46,380 41,914 72 0 43 2 0 3,744
2026-09-23 40,040 35,066 63 0 10 3 0 3,974
2026-09-22 50,943 39,204 4,187 0 3,200 12 0 3,646
2026-09-21 18,485 13,372 948 0 385 1 0 3,589

@JaeYeonLee0621 JaeYeonLee0621 self-assigned this Sep 29, 2026
request_log.processing_time_ms = int((response_start_time - kwargs["request_start"]) * 1000)
request_log.response_time_ms = int((end_time - response_start_time) * 1000)
request_log.status_code = result.status_code

@JaeYeonLee0621 JaeYeonLee0621 Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • Not streaming (JsonResponse) : log_request saves result.status_code as-is
  • Streaming : log_request leaves it None, and _openai_stream figures out the right code when the stream actually ends

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe I am misunderstanding but should we not set status_code to 499 here if it is None?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Or can non-streaming requests never lead to a disconnect this way? What happens if a client disconnects in a non streaming request before the result is returned? I guess we still wait for a response and then just take that status code? But I guess the whole request in Django is then closed so do we even listen to the response then?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I may be wrong, but I'm not sure if client disconnects are handled at all for non-streaming responses. The docs only mention it for StreamingHttpResponse. If I were to guess, I'd say you'd probably get a broken pipe error or something when Django finishes processing and tries to write the response - I imagine this would show up in the logs though? But that's just a guess on my side.

@JaeYeonLee0621 JaeYeonLee0621 Oct 5, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't know about this one that much, but I studied with deepseek 🐳 for a while, as far as I understand :

  • In Django's ASGIHandler, when a client disconnects while the view is still awaiting the upstream.
  • It cancels the request task, which raises asyncio.CancelledError inside the view. (not a broken pipe, because the task is cancelled before any write attempt)
  • The except asyncio.CancelledError catches it and records 499 for non-streaming responses too.


request_log.status_code = 200
yield "data: [DONE]\n\n"
except (asyncio.CancelledError, GeneratorExit):

@JaeYeonLee0621 JaeYeonLee0621 Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • GeneratorExit: Signals a paused generator to immediately stop producing data and clean up because the consumer called .close().
  • asyncio.CancelledError: Signals a paused coroutine or async task to abort immediately because the task was cancelled.
  • Why Exception cannot catch: Both inherit directly from BaseException rather than Exception.

@JaeYeonLee0621
JaeYeonLee0621 requested review from meffmadd and natkam and removed request for natkam September 29, 2026 13:43
@JaeYeonLee0621
JaeYeonLee0621 marked this pull request as ready for review September 29, 2026 13:43
return Usage(input_tokens=0, output_tokens=0)


def _openai_stream(

@JaeYeonLee0621 JaeYeonLee0621 Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • catch_router_exceptions : runs at the very start, only ever report upstream errors
  • _openai_stream : runs during streaming, by the time _openai_stream is running, both things can happen there: the client closing, and the upstream failing mid-stream

@JaeYeonLee0621 JaeYeonLee0621 linked an issue Sep 29, 2026 that may be closed by this pull request

@meffmadd meffmadd left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks very good! I have one quick question for how non-streaming requests should be handled.

request_log.processing_time_ms = int((response_start_time - kwargs["request_start"]) * 1000)
request_log.response_time_ms = int((end_time - response_start_time) * 1000)
request_log.status_code = result.status_code

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe I am misunderstanding but should we not set status_code to 499 here if it is None?

request_log.processing_time_ms = int((response_start_time - kwargs["request_start"]) * 1000)
request_log.response_time_ms = int((end_time - response_start_time) * 1000)
request_log.status_code = result.status_code

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Or can non-streaming requests never lead to a disconnect this way? What happens if a client disconnects in a non streaming request before the result is returned? I guess we still wait for a response and then just take that status code? But I guess the whole request in Django is then closed so do we even listen to the response then?


response_start_time = time.monotonic()
result: HttpResponse | StreamingHttpResponse = await view_func(request, *args, **kwargs)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I add try, except for non-streaming requests to save 499 status code and stop running Django logic.

@natkam natkam left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd just remove that if __name__ = ... in the test file, but generally it looks good to me!

Comment thread aqueduct/gateway/tests/test_stream_status.py Outdated
request_log.processing_time_ms = int((response_start_time - kwargs["request_start"]) * 1000)
request_log.response_time_ms = int((end_time - response_start_time) * 1000)
request_log.status_code = result.status_code

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I may be wrong, but I'm not sure if client disconnects are handled at all for non-streaming responses. The docs only mention it for StreamingHttpResponse. If I were to guess, I'd say you'd probably get a broken pipe error or something when Django finishes processing and tries to write the response - I imagine this would show up in the logs though? But that's just a guess on my side.

Comment thread aqueduct/gateway/tests/test_stream_status.py Outdated

@natkam natkam left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! I'll merge it :)
EDIT: should I merge it? (sorry!)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Random non-responding requests invisible in the failure stats

3 participants