You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Completed Pi reasoning can render as Thought for 0ms even when the assistant response took several seconds. BB loses the response start time, batches the reasoning lifecycle events, gives the full batch one storage timestamp, and then subtracts equal start and completion timestamps.
Versions and environment
bb CLI 0.42.1 reading desktop-app thread data; reported desktop build is from the personal fork
Current upstream main: dba32a469fd820ff6106715db0aaf6ed297d79a5
macOS 15.7.7, Node.js 22.23.1
Pi 0.85.1
BB provider: pi
Source provider/model: openai-codex/gpt-5.6-sol, reasoning level high
Worktree environment
Steps to reproduce
Start a Pi thread with openai-codex/gpt-5.6-sol and high reasoning.
Run a tool-use task that produces several short visible reasoning summaries between tool calls.
Wait for the turn to complete.
Run bb thread log <thread-id> --all --format verbose.
Inspect the completed thought titles.
Captured repro thread: thr_7unpk42ev3. I did not create a public replacement thread because the captured raw and normalized events show the full timing path.
Expected vs actual
Expected: A truthful positive duration, or `Thought` if BB has no reliable duration.
Actual:
── Thought for 0ms
**Searching code for Braintrust API**
Evidence
For reasoning item da5a9e143a-i106, the normalized BB events are:
Across the captured thread, 637 of 1,453 completed reasoning items have an exact zero duration (43.8%). I measured this by pairing each persisted reasoning item/started with its item/completed event and subtracting their createdAt values. The first captured zero is from 2026-08-21, before PR #3066 made completed thoughts visible.
Observed code path on upstream main:
Pi accepts message_start only for custom messages, so an assistant boundary and its source timestamp are ignored:
No relevant errors appear in ~/.bb/logs; those logs show normal timeline builds and provider-session cleanup for this thread.
Suggested fix
As the smallest safe display fix, render Thought instead of Thought for 0ms when completedAt <= startedAt.
To preserve a real duration, carry event occurrence time through the provider bridge and host-daemon event envelope, then persist each event's own time instead of one database time for the full batch. For Pi reasoning returned as a completion-time summary, use the assistant message_start timestamp as the start.
What you ruled out
This is not a duration formatter bug: it correctly receives and formats a zero difference.
Searches of open and closed issues for Thought for 0ms, 0ms thinking, thinking duration, and reasoning timestamp found no matching report.
The relevant files are identical in the personal fork and upstream main at dba32a469fd820ff6106715db0aaf6ed297d79a5.
Suggested priority and effort
Medium priority. It makes a visible timing claim wrong for many Pi/OpenAI Codex thoughts. Omitting an unreliable duration is a small workaround; preserving the real duration needs a cross-contract timing change.
Summary
Completed Pi reasoning can render as
Thought for 0mseven when the assistant response took several seconds. BB loses the response start time, batches the reasoning lifecycle events, gives the full batch one storage timestamp, and then subtracts equal start and completion timestamps.Versions and environment
main:dba32a469fd820ff6106715db0aaf6ed297d79a5piopenai-codex/gpt-5.6-sol, reasoning levelhighSteps to reproduce
openai-codex/gpt-5.6-soland high reasoning.bb thread log <thread-id> --all --format verbose.Captured repro thread:
thr_7unpk42ev3. I did not create a public replacement thread because the captured raw and normalized events show the full timing path.Expected vs actual
Evidence
For reasoning item
da5a9e143a-i106, the normalized BB events are:The corresponding raw Pi assistant message has:
Across the captured thread, 637 of 1,453 completed reasoning items have an exact zero duration (43.8%). I measured this by pairing each persisted reasoning
item/startedwith itsitem/completedevent and subtracting theircreatedAtvalues. The first captured zero is from 2026-08-21, before PR #3066 made completed thoughts visible.Observed code path on upstream
main:message_startonly for custom messages, so an assistant boundary and its source timestamp are ignored:bb/plugins/provider-pi/src/delta-translation.ts
Lines 109 to 120 in dba32a4
bb/plugins/provider-pi/src/delta-translation.ts
Lines 565 to 580 in dba32a4
thinking_deltaandthinking_endwithout source time:bb/plugins/provider-pi/src/delta-translation.ts
Lines 740 to 795 in dba32a4
bb/apps/host-daemon/src/event-sink.ts
Line 11 in dba32a4
bb/apps/host-daemon/src/event-sink.ts
Lines 214 to 261 in dba32a4
bb/packages/host-daemon-contract/src/session.ts
Lines 189 to 197 in dba32a4
Date.now()value for every accepted event in the batch:bb/packages/db/src/data/events.ts
Lines 663 to 718 in dba32a4
bb/packages/thread-view/src/reasoning-lifecycle-projection.ts
Lines 177 to 195 in dba32a4
bb/packages/thread-view/src/reasoning-lifecycle-projection.ts
Lines 217 to 237 in dba32a4
No relevant errors appear in
~/.bb/logs; those logs show normal timeline builds and provider-session cleanup for this thread.Suggested fix
As the smallest safe display fix, render
Thoughtinstead ofThought for 0mswhencompletedAt <= startedAt.To preserve a real duration, carry event occurrence time through the provider bridge and host-daemon event envelope, then persist each event's own time instead of one database time for the full batch. For Pi reasoning returned as a completion-time summary, use the assistant
message_starttimestamp as the start.What you ruled out
Thought for 0ms,0ms thinking,thinking duration, andreasoning timestampfound no matching report.mainatdba32a469fd820ff6106715db0aaf6ed297d79a5.Suggested priority and effort
Medium priority. It makes a visible timing claim wrong for many Pi/OpenAI Codex thoughts. Omitting an unreliable duration is a small workaround; preserving the real duration needs a cross-contract timing change.