DEV Community

Dakota Lin
Dakota Lin

Posted on

I Logged Three Clocks. The Model Sat Still.

I blamed the model before I split the clocks. My first blame was lazy and flat wrong. The queue and the wire owned the wait.

You feel a coding agent stall and you swear. You swear at weights you never even measured. Have you timed the three gaps, or just the mood?

I mean three clocks, not one sad total. I log send time, first byte, and done. If those three stay glued, your graph is a rumor.

I wrote a tiny harness and refused the folklore. Same prompt, same bytes, and two endpoints. The only story I trust is a span I can replay.

I point one run at MonkeyCode's free model access. Disclosure: This article was prepared as part of MonkeyCode's product outreach. Their free server is a hop, not a trophy.

I do not print a quota, a chip, or a lease. Those facts move fast, and stale numbers lie. You should read the live page before you quote one.

I would keep the other run on a box I already own. That contrast is the whole experiment, nothing fancier. If both hops look alike, the model was never the villain.

Here is the harness I keep beside the notes. Treat the JSON body as a local stand-in, not their API. Adapt every field to docs you verified today.

#!/usr/bin/env python3
"""Three-clock sketch. Unexecuted example. No live timings claimed."""
import json, os, time, urllib.request

ENDPOINT = os.environ["MODEL_ENDPOINT"]
TOKEN = os.environ.get("MODEL_TOKEN", "")
PROMPT = open("fixture.txt", encoding="utf-8").read()

def one_shot(body: bytes) -> dict:
    req = urllib.request.Request(ENDPOINT, data=body, method="POST")
    req.add_header("Content-Type", "application/json")
    if TOKEN:
        req.add_header("Authorization", "Bearer " + TOKEN)
    t_send = time.perf_counter()
    with urllib.request.urlopen(req, timeout=12) as resp:
        t_headers = time.perf_counter()
        raw = resp.read()
        t_done = time.perf_counter()
    return {
        "to_headers_ms": round((t_headers - t_send) * 1000, 1),
        "body_ms": round((t_done - t_headers) * 1000, 1),
        "total_ms": round((t_done - t_send) * 1000, 1),
        "bytes": len(raw),
    }

if __name__ == "__main__":
    payload = json.dumps({"prompt": PROMPT}).encode()
    for i in range(5):
        row = one_shot(payload)
        row["i"] = i
        row["cold"] = i == 0
        print(json.dumps(row))
        time.sleep(1.0)
Enter fullscreen mode Exit fullscreen mode

Run it only against an endpoint you may call. Keep the token in the environment, never in git. Five loops are a sketch, not a promise to the wire.

export MODEL_ENDPOINT="https://example.invalid/complete"
export MODEL_TOKEN="set-me-outside-the-repo"
python3 three_clocks.py | tee rows.jsonl
Enter fullscreen mode Exit fullscreen mode

I tag each row warm or cold before I plot it. A cold socket will impersonate a slow model. Have you been graphing a handshake and calling it smart?

Header time on a buffered call is not token time. The server may hold every byte until the end. If you need the first token, you must force a stream.

So the second script reads small chunks on purpose. It stamps the clock when the first chunk lands. The rest of the bar is whatever remains until close.

def stream_gaps(body: bytes) -> dict:
    req = urllib.request.Request(ENDPOINT, data=body, method="POST")
    req.add_header("Content-Type", "application/json")
    if TOKEN:
        req.add_header("Authorization", "Bearer " + TOKEN)
    t_send = time.perf_counter()
    t_first = None
    with urllib.request.urlopen(req, timeout=12) as resp:
        while True:
            chunk = resp.read(64)
            if not chunk:
                break
            if t_first is None:
                t_first = time.perf_counter()
    t_done = time.perf_counter()
    if t_first is None:
        t_first = t_done
    return {
        "to_first_ms": round((t_first - t_send) * 1000, 1),
        "after_ms": round((t_done - t_first) * 1000, 1),
    }
Enter fullscreen mode Exit fullscreen mode

If the server never flushes, those two gaps collapse. You just measured a buffered body with extra steps. Check the docs for a real stream flag before you blame the hop.

I print an ASCII bar and I refuse a theme. Two bars per row, first chunk and the rest. The picture is a receipt, not a ranking post.

def bar(ms: float) -> str:
    cells = int(ms // 25)  # 25 ms per mark, then stop the ink
    return "#" * max(0, min(cells, 40))

def render(rows: list) -> str:
    lines = []
    for row in rows:
        lines.append("%s first %s %s" % (row["i"], bar(row["to_first_ms"]), row["to_first_ms"]))
        lines.append("%s rest  %s %s" % (row["i"], bar(row["after_ms"]), row["after_ms"]))
    return "\n".join(lines)
Enter fullscreen mode Exit fullscreen mode

Look at a row before you touch the prompt. If the first bar dwarfs the rest, the model sat still. Why rewrite instructions while the queue eats the span?

Picture a diner with a fast cook and a slow bell. You yell at the kitchen because the plate is late. The ticket was stuck on the rail the whole time.

That is this graph, with worse coffee and more JSON. The cook is the work after the first chunk. The rail is queue, DNS, TLS, and your own retry nap.

I keep the graph that shows the rail, not the cook. A teammate can argue with a bar of hashes. A teammate cannot argue with your mood from Tuesday.

When the rail is long, I pull the call off the key path. Batch the review after lunch, or night the notes. Do not park a remote hop inside a save hook.

A free server is a fair night shift for that batch. It is a bad stand-in for a compiler. Would you block one keystroke on a queue you do not own?

I still want that hop for work that can wait. A diff read after lunch can wait a little. Autocomplete under a finger cannot, and should not pretend.

One request in flight is the whole load I allow. A parallel fan-out would only measure my stampede. Shared free capacity is a commons, not a drum I get to hit.

Secrets do not ride a hop I do not operate. If the fixture holds private source, I strip it first. A friendly free box is still someone else's machine.

I also refuse one evening as a service level. Other people arrive later, and the rail grows. Your pretty bar from tonight can lie by morning.

Shared ids beat two clocks that never met. Put one request id on the client row and the log. Without that id, you compared neighbors, not one trip.

Here is the gate I run before I trust a bar. Same id, same byte count, and a cold flag. If any one is missing, I throw that row out.

def keep(row: dict) -> bool:
    needed = ("req_id", "bytes", "cold", "to_first_ms", "after_ms")
    return all(k in row and row[k] is not None for k in needed)
Enter fullscreen mode Exit fullscreen mode

Who should skip this approach, and stay skipped? Skip it when policy forbids a third-party model hop. Skip it when you need a model menu I will not invent.

Skip it when your budget is one human keystroke. Skip it when you wanted a winner and a medal. I kept a method, not a ranking, and that was the point.

What should you change after one honest graph? Change the path, the batching, or the timeout. Do not change the prompt until the rail gets short.

I set the client timeout to a number I can say aloud. Twelve seconds makes a stuck queue visible and boring. Sixty seconds hides that same queue inside a spinner.

Retries are where a clean graph goes to die. A hidden second try adds a full extra rail. Log the attempt index, or the total will confess to fiction.

A total can include two naps and still look sincere. The model may have finished on the first try. Which clock did you paste into the incident note?

DNS and TLS deserve blame on a fresh process. Reuse the connection when the API actually allows it. Otherwise every row opens with a handshake costume party.

I store a warm series beside a separate cold row. Mixing them is how a free hop gets framed. The costume party is not the model's personality.

Payload size will fake a generation problem too. A novel-length diff is a shipping choice you made. Trim the fixture before you indict the remote side.

I keep the fixture in the repo next to the script. A teammate should rerun the same bytes, not a vibe. If you edit the fixture, you started a different experiment.

The graph stays ugly so I cannot decorate a miss. No badge and no theme, and no weekday duel. Two bars, a request id, and a cold flag are enough.

If the rest bar grows and the first stays flat, I open the prompt. That is the only time the model owes me an answer. Until then it can sit still, and it should.

Free model access does not mean a queue-free lane. A shared server spends time on other callers. Treat that wait as traffic, not as a broken file of weights.

I will not tell you the wait is small or large. I did not run this harness for a published score. Print your own rows, then keep the graph that matches them.

Check the project page on the day you actually run. Names, limits, and hardware are their facts, not mine. If this note and that page disagree, believe the page.

Point that sketch at their free server when the wait fits. Read the live terms, then keep or drop the endpoint. I would rather you discard it than quote a number I never printed.

Top comments (0)