正在加载评测记录…
本页仅复测 Tube,开启跨轮 memory;max_turns=80,使用同一 seed / layout。
CROSS-RUN MEMORY
这一轮实际使用的经验
日志确认 GPT 在首次运动前读取了 3 个文件,并成功写回草稿 11 次。下方保留启动快照和本轮草稿;历史经验不替代原生判定。这轮还改进了搬运后的抓取复核,不能把变化单独归因于 memory。
MEMORY · 已读取
# Layered memory index Generated from memory leaf frontmatter. ## Global - [Key: plate-normal grasp and rejected handover poses](global/insert_key_gpt_only_observations.md) — insert_key, GPT-only physicalrsi_v7 through v11 - [Tubes: placement evidence, path ordering and tracking correction](global/insert_tubes_gpt_only_observations.md) — insert_tubes, GPT-only physicalrsi_v11 stock test_tube/00002 - [Latest tube transport outcome and remaining grasp risk](global/insert_tubes_transport_grasp_v11.md) — insert_tubes; final v11 evidence for v12 memory iteration - [Evidence and recovery rules for coded insertion](global/insertion_evidence_rules.md) — GPT-only staged insert_key or insert_tubes, stock RoboDojo assets ## Suite _(none)_
insert_key_gpt_only_observations · 快照中保留,未观察到读取
--- id: insert_key_gpt_only_observations scope: global kind: failure title: 'Key: plate-normal grasp and rejected handover poses' applies_when: insert_key, GPT-only physicalrsi_v7 through v11 confidence: single-shot evidence: cells: - v6_insert_key_s0_l0 - v7_insert_key_s0_l0 - v11_insert_key_s0_l0 attempts: 3 reviewed: '2026-09-29' provenance: Native results, stage traces and videos from the insertion campaign; human-reviewed synthesis. --- ## Verified observations, limited to this layout The key plate is local X-Z and its normal is local Y. Closing on the circular bow edges along local X caused a drop after giver release in v6. A plate-normal receiving grasp remains the intended geometry; it has NOT yet been demonstrated in native simulation. In v7 the required roll was rejected at the initial right TCP. In v11 moving the open right hand near [0.160,-0.208,0.926] succeeded, but BOTH equivalent jaw rolls were rejected by IK. This is negative evidence for the current fixed presentation/staging geometry, not proof that the task is mechanically impossible. About 10.4 degrees of key settling occurred in the giver during receiver positioning. The controller now allows up to 20 degrees during receiving stages while preserving the original grasp reference, the 8 mm translation guard, measured jaw-normal checks, and independent lift verification after giver release. ## Current limitation / next hypothesis The current template exposes no free handover pose parameter. Memory cannot authorize unavailable movement tools or force a rejected pose. A future controller may need a different giver presentation pose and a paired reachability check for both arms before contact. Do not claim that re-reading memory or switching a hole fixes this key kinematic limit. For insertion only after VERIFIED transfer: key tip local -Z points down, directed tooth side local +X matches slot local +X before the final turn. Recompute geometry from live poses; never replay old XYZ coordinates as targets.
insert_tubes_gpt_only_observations · 已读取
--- id: insert_tubes_gpt_only_observations scope: global kind: strategy title: 'Tubes: placement evidence, path ordering and tracking correction' applies_when: insert_tubes, GPT-only physicalrsi_v11 stock test_tube/00002 confidence: single-shot evidence: cells: - v6_insert_tubes_s0_l0 - v9_insert_tubes_s0_l0 - v11_insert_tubes_s0_l0 attempts: 3 reviewed: '2026-09-29' provenance: Native results, stage traces and videos from the insertion campaign; human-reviewed synthesis. --- ## Observed Grasp the upper body below the cap and verify a lift. Tube0 and tube1 completed placement cycles in v6. Tube2 was verified upright, but crossed the occupied right-end tube while approaching a central hole; it tilted about 59 degrees and lost the grasp. Collision is the leading video/geometry explanation, without direct contact telemetry. v9 raised above all other tubes, even off-route ones, and accumulated 3-4 mm execution bias during alignment. Repeating the same target did not remove that bias and consumed the native budget. v11 filters obstacles by the transport route and uses measured endpoint corrections. In the current v11 pilot, the first two tubes were released and stable by native step 308 without failed-stage recovery; the third is still under evaluation. A placed tube may lean while satisfying fresh native conditions. Do not override fresh native containment/depth/upright evidence solely because root XY differs from the chosen hole center. Homing is deferred until all three tubes are released, then call the last tube's home stage again until both used arms are home. ## Untested next planning hypothesis For this layout, try tube0 into central hole 2 first, tube1 into left outer hole 0, and tube2 into right outer hole 4. This may avoid carrying the final right-hand tube across an already occupied right-end hole and reduce required lifting/travel. This ordering is NOT a proven successful recipe. Inspect fresh locations and occupancy before choosing it; controller guards remain authoritative. Hole rows and coordinates are rack-relative, not fixed world poses. ## Record For each tube retain object id, chosen hole, verified lift, alignment error, insertion/release evidence, and home status. Three inserted tubes without required arm home conditions still fail the full native task.
insert_tubes_transport_grasp_v11 · 已读取
--- id: insert_tubes_transport_grasp_v11 scope: global kind: failure title: Latest tube transport outcome and remaining grasp risk applies_when: insert_tubes; final v11 evidence for v12 memory iteration confidence: single-shot evidence: cells: - v11_insert_tubes_s0_l0 attempts: 1 --- ## Final update to the v11 pilot This supersedes the earlier note that tube2 was still under evaluation. The final native step was 443. Tube0 and tube1 were released and stable. Tube2 reached the high transport path, then about 9.76 mm translation and 5.42 degrees of relative grasp change tripped the 8 mm guard; it was no longer considered held. Its upright pose could not use the lying-on-table pickup template. The operator stopped unchanged retries. The task did not succeed and was an interrupted pilot, not a complete benchmark failure. ## v12 controller change to verify, not assume Moderate settling (at most 20 mm / 20 degrees) during clear transport now requires a fresh independent lift before accepting a new grasp reference, as already done after rotation. Loss of following still blocks the stage. This has offline regression coverage but no native success evidence yet. Do not manually relax the insertion or grasp guards. The center-first hole ordering in the prior tube strategy note remains an untested planning hypothesis. Reusing this memory does not itself establish improvement; inspect the actual choices, grasp evidence, native step use and whole-task verdict.
insertion_evidence_rules · 快照中保留,未观察到读取
--- id: insertion_evidence_rules scope: global kind: strategy title: Evidence and recovery rules for coded insertion applies_when: GPT-only staged insert_key or insert_tubes, stock RoboDojo assets confidence: single-shot evidence: cells: - v6_insert_key_s0_l0 - v6_insert_tubes_s0_l0 attempts: 2 reviewed: '2026-09-29' provenance: Native results, stage traces and videos from the insertion campaign; human-reviewed synthesis. --- ## Scope and evidence All observations are seed 0 / layout 0 native pilots. No full-task success has been observed. Some runs were operator interrupted, so they are not complete benchmark failures. Only native eval_success establishes whole-task success. ## Apply Read the current controller phase and fresh object poses. Track separate verified facts and hypotheses in STATE. A successful pose plan is not contact evidence. Gripper state is the previous control command, not measured finger contact. Object following during an independent lift is required to verify grasp/transfer. Do not repeat a failed unchanged stage. Changing hole_index does not change a pickup pose, so it cannot repair the same unreachable pickup. A bounded recover is useful only while a grasp remains verified. Keep the original 300/500 native limits and reserve enough actions for required finishing/homing. Published memory is read-only. Write a short episode draft after a new milestone or failure; include native step and observed evidence. Do not create a successful recipe for partial placements.
GPT 写回的本轮经验草稿
这是待核验草稿;结论需与原生结果对照。
Insert tubes physicalrsi_v12 step397;103 native steps remain. Observed tubes0/1 released holes2/0 with native geometry passing. Tube2 right verified held sideways87.8deg. Recover selecting hole9 changed cap-up target to [.5012,-.1737,.966],quat[.2338,-.0188,.0048,.9721], but plan failed0steps, probe97%, orientation blocker. Hole4 also rejected previously; lift recovery alone insufficient; transport/handover skips and restart verified-held rejected. Hypothesis: another free rack aperture changes desired yaw sufficiently for planning. No overall success.
对照原生记录后合并的本轮经验
下轮可以读取;本轮启动时尚不存在,不属于本轮输入。
--- id: insert_tubes_memory_v12_outcome scope: global kind: failure title: 'Latest v12: memory used; two placements, third orientation rejected' applies_when: insert_tubes, GPT-only physicalrsi_v12, seed 0 / layout 0; read with earlier observations confidence: single-shot evidence: cells: - v12_insert_tubes_s0_l0 attempts: 1 reviewed: '2026-09-29' provenance: Reviewed against result.json, states.json, run.log, GPT draft and native video; not automatic promotion of model prose. --- ## Outcome and scope This supersedes any assumption that the center-first order is a validated complete recipe. The v12 pilot read MEMORY.md and two tube leaves before its first motion, chose tube0/hole2, tube1/hole0, tube2/hole4, and successfully updated its episode draft 11 times. It used 52 GPT turns and 397 native steps, then was deliberately operator-interrupted after repeated unreachable orientation plans. Raw termination_reason is infrastructure_error from the explicit operator stop, not an upstream API outage. It is an invalid/interrupted trial, not a complete benchmark failure or a success. ## Verified partial progress Tube0 completed release, retreat and placement checks at native step 180 in central hole2. Its transport required independent grasp re-verification (22.32 mm lift), then a bounded recovery after the above-hole endpoint check failed. Tube1 completed in left hole0 at step332 without blocked-stage recovery. At final step397 both depth predicates were current and true; containment was last evaluated true at step388, and fresh axis measurements were about 0.38 and 7.29 degrees. Do not relabel the older containment evaluations as current-step evaluations. Tube2 was grasped by the right hand and independently lifted 22.41 mm at step387, but remained about 87.8 degrees from upright. Its cap-up target was rejected for orientation at hole4, again after a recovery lift, and again after selecting hole9. All three failed orientation plans executed zero native actions. The other two tubes remained released. Neither arm met the required home condition; native eval_success remained false. ## Failed approaches and limits Changing height with recover did not make this third-tube orientation target reachable. Changing hole4 to hole9 altered the yaw but also failed. Transport and handover could not bypass the required orient stage; restart was correctly blocked while the grasp remained verified. Do not use those rejected stage changes as a recovery recipe. The generic arm-change error mentions handover, but this tube phase does not expose a usable handover path. The draft's suggestion to try another aperture is an untested hypothesis, not evidence that it would fix reachability. The intended final hole4 transport was never reached, so this trial does not establish whether the center-first order avoids the historical last-tube collision. Memory and transport re-verification code both changed in v12; this is not a memory-only ablation. ## Next controller hypothesis, not an executable recipe Evaluate a supported held-tube translation toward a reachable orientation staging pose before rotating, or choose a different verified pickup transform. Check the candidate orientation before committing contact motion and retain lift verification and stage guards. No such change was tested in v12; memory cannot make an unavailable controller action callable. Budget native steps for all three placements plus both arms returning home. Only fresh native success can promote a full task recipe.
PROTOCOL
这次测了什么
规划与执行
GPT-6 Astra,medium reasoning,经 OpenRouter Responses API;RPent lean harness,最近两轮历史、STATE 记录与最新观测。GPT 调用 insertion_stage,代码执行抓取、换手、摆正、对孔、分段插入和撤离。policy_actions 为 0;原始 move_to / reset 工具在此模式下已替换。
输入与成功判定
头部和腕部图像、物体真实位姿与原生条件反馈均对 GPT 可见,包含特权状态信息。只有原生 eval_success=true 才算成功;单个深度或朝向条件通过不代表任务完成。
测试边界
204 上的原生仿真环境,非官方 Docker;key 300 步,tubes 500 步;每项最多 80 个 GPT 回合,与提示词基线的 200 回合预算不同。渲染有至少一帧延迟警告。每项只测 seed 0 / layout 0,不能由此推断 42 个任务总分或稳定成功率。
规则来源
此前成功与失败轨迹,以及 PhysicalRSI 动作视频的观察。动作顺序参考视频,尺寸按本地原版资产标定;这些是待验证的控制模板。视频不能精确证明松手深度,也不能证明本控制器会成功。PhysicalRSI 项目。
RPent 22d104d3 · 控制器 physicalrsi_v12 · 下载本页结果与操作记录 JSON