feat: operate ChatGPT's new page layout, and fix the related defects #79

Merged
xicv merged 34 commits from feat/new-layout-port into main 2026-09-25 03:37:05 +00:00
xicv commented 2026-09-25 03:28:39 +00:00 (Migrated from github.com)

Summary

Fixes #77 on top of the fail-fast step (#78). Ego Chat now operates ChatGPT's new page layout, and this PR also fixes the related defects found while diagnosing it.

ChatGPT's new page layout. The driver now operates both layouts. Details are live-read on 2026-09-24/25.

  • Composer: a ProseMirror [role="textbox"] inside [data-composer-layout]. The textarea it shows for about two seconds while hydrating counts as still loading.
  • Send and Stop: the form's submit button.
  • Messages: turns are [data-turn-key], and each message unit is split by role.
  • Composer surface: [data-thread-scroll-footer].
  • Model menu: the "Select ChatGPT model" trigger, which closes on focus.
  • The new layout's head is tail-v2. It matches a tail-v1 head across layouts by message ID, assistant role and non-blank text.
  • The answering model is read from ChatGPT's own conversation data. The driver first looks among the page's buffered responses, then makes one audited plain read in the page that returns only the slug.
  • Attachment capture stops as attachment_capture_unsupported_layout.

Pro ladder, in the user's order: GPT-6 Pro, then Latest at Extra High (latest_extra_high, effort state extra_high_fallback), then GPT-5.6 Sol Pro, then a stop. The planned rung's answer (gpt-6 / gpt-5-6-thinking) is silent in downgrade notifications, and a pause convergence accepts it.

Nothing sent, said plainly. Every exchange carries delivery.state (not_sent / confirmed / unknown) and one delivery.nextAction, and supervision reports the send state. One page reason repeating for ten minutes stops as pre_send_stalled instead of retrying for 90 minutes. SKILL.md and the MCP instructions tell hosts what to do for each action.

Related defects:

  • Policy reload. model-policy.json is read again whenever it changes. broker-status shows loadedAt and stale.
  • Vanished Spaces. ego_ensure_model_policy and ui-check used to stop as bound_task_space_identity_changed after every Ego Lite restart. They now recover the Space through verify and try once more. Their caller timeouts cover all three runs.
  • Lost create-once tab. An unbound create-once binding that lost its tab or Space is re-pinned to a blank starting page in its own Space, with a new repin_unbound mode that never composes. The binding must have recorded no Send. No other exchange on it may be running or may have sent, although abandoned attempts with no Send evidence are accepted. Otherwise the exchange stops at once as create_once_repin_blocked or create_once_repin_failed, both proven pre-Send stops.
  • Project URL forms and the tab leak (found live today). ChatGPT redirects a project conversation between g-p-<id> and g-p-<id>-<title>. A binding stored in the plain form treated its own page as another chat and opened a new tab on every retry: 128 tabs, about 23 GB in Ego Lite. The driver now identifies a conversation by project and conversation ID, reports the recorded form, and reuses an open tab of the conversation.

Verification

  • npm test: 1522 tests, 1521 pass, 1 skipped. npm run test:ego-monitor 148/148. npm run lint clean. A3K fixture regenerated.
  • New tests were seen failing first, or were mutation-checked when the code came first: the re-pin driver's claimed-tab filter, project URL identity, open-tab reuse.
  • Live on a dev broker (2026-09-25):
    • Exchanges on the new layout succeeded at latest_extra_high with a tail-v2 head.
    • The first exchange left responseModelSlug null because the page's response was gone. That led to the page-read fallback. The next exchange recorded gpt-5-6-thinking, with no downgrade alert.
    • No tabs leaked.
  • Codex review (read-only) of the whole branch found four P2s, all fixed with tests that fail without the fix:
    • the Extra High rung now holds only on the route named Latest (LATEST_ROUTE_LABELS);
    • after three failed re-pin runs the exchange stops as create_once_repin_failed instead of retrying;
    • a new-layout answer that is only a picture is image-only again;
    • a convergence stopped on the review it just captured names it in reviewStopWorkflowId, so supervision still reports it as sent.
  • Codex re-review of the fixes found one more: the streak stagnation and cycle-limit stops needed the same marker. Fixed with tests; a final Codex review found no actionable regression. (Codex's read-only sandbox cannot run the suites; they were run locally.)

Residual risk, stated in TECH.md: if an abandoned attempt did deliver without the broker seeing it, re-pinning its binding starts a second conversation.

## Summary Fixes #77 on top of the fail-fast step (#78). Ego Chat now operates ChatGPT's new page layout, and this PR also fixes the related defects found while diagnosing it. **ChatGPT's new page layout.** The driver now operates both layouts. Details are live-read on 2026-09-24/25. - Composer: a ProseMirror `[role="textbox"]` inside `[data-composer-layout]`. The textarea it shows for about two seconds while hydrating counts as still loading. - Send and Stop: the form's submit button. - Messages: turns are `[data-turn-key]`, and each message unit is split by role. - Composer surface: `[data-thread-scroll-footer]`. - Model menu: the "Select ChatGPT model" trigger, which closes on focus. - The new layout's head is `tail-v2`. It matches a `tail-v1` head across layouts by message ID, assistant role and non-blank text. - The answering model is read from ChatGPT's own conversation data. The driver first looks among the page's buffered responses, then makes one audited plain read in the page that returns only the slug. - Attachment capture stops as `attachment_capture_unsupported_layout`. **Pro ladder, in the user's order:** GPT-6 Pro, then Latest at Extra High (`latest_extra_high`, effort state `extra_high_fallback`), then GPT-5.6 Sol Pro, then a stop. The planned rung's answer (`gpt-6` / `gpt-5-6-thinking`) is silent in downgrade notifications, and a pause convergence accepts it. **Nothing sent, said plainly.** Every exchange carries `delivery.state` (`not_sent` / `confirmed` / `unknown`) and one `delivery.nextAction`, and supervision reports the send state. One page reason repeating for ten minutes stops as `pre_send_stalled` instead of retrying for 90 minutes. SKILL.md and the MCP instructions tell hosts what to do for each action. **Related defects:** - **Policy reload.** `model-policy.json` is read again whenever it changes. `broker-status` shows `loadedAt` and `stale`. - **Vanished Spaces.** `ego_ensure_model_policy` and `ui-check` used to stop as `bound_task_space_identity_changed` after every Ego Lite restart. They now recover the Space through verify and try once more. Their caller timeouts cover all three runs. - **Lost create-once tab.** An unbound create-once binding that lost its tab or Space is re-pinned to a blank starting page in its own Space, with a new `repin_unbound` mode that never composes. The binding must have recorded no Send. No other exchange on it may be running or may have sent, although abandoned attempts with no Send evidence are accepted. Otherwise the exchange stops at once as `create_once_repin_blocked` or `create_once_repin_failed`, both proven pre-Send stops. - **Project URL forms and the tab leak (found live today).** ChatGPT redirects a project conversation between `g-p-<id>` and `g-p-<id>-<title>`. A binding stored in the plain form treated its own page as another chat and opened a new tab on every retry: 128 tabs, about 23 GB in Ego Lite. The driver now identifies a conversation by project and conversation ID, reports the recorded form, and reuses an open tab of the conversation. ## Verification - `npm test`: 1522 tests, 1521 pass, 1 skipped. `npm run test:ego-monitor` 148/148. `npm run lint` clean. A3K fixture regenerated. - New tests were seen failing first, or were mutation-checked when the code came first: the re-pin driver's claimed-tab filter, project URL identity, open-tab reuse. - Live on a dev broker (2026-09-25): - Exchanges on the new layout succeeded at `latest_extra_high` with a `tail-v2` head. - The first exchange left `responseModelSlug` null because the page's response was gone. That led to the page-read fallback. The next exchange recorded `gpt-5-6-thinking`, with no downgrade alert. - No tabs leaked. - Codex review (read-only) of the whole branch found four P2s, all fixed with tests that fail without the fix: - the Extra High rung now holds only on the route named Latest (`LATEST_ROUTE_LABELS`); - after three failed re-pin runs the exchange stops as `create_once_repin_failed` instead of retrying; - a new-layout answer that is only a picture is image-only again; - a convergence stopped on the review it just captured names it in `reviewStopWorkflowId`, so supervision still reports it as sent. - Codex re-review of the fixes found one more: the streak stagnation and cycle-limit stops needed the same marker. Fixed with tests; a final Codex review found no actionable regression. (Codex's read-only sandbox cannot run the suites; they were run locally.) Residual risk, stated in TECH.md: if an abandoned attempt did deliver without the broker seeing it, re-pinning its binding starts a second conversation.
Sign in to join this conversation.
No description provided.