Letting AI Run Research: Which Decisions Should I Delegate?
让 AI 自主推进科研:我在尝试委托哪些判断?
Read and discuss with AI · Open in Claude / ChatGPT
Continue discussing the article with AI. If the prompt is missing after opening, copy it below.
As I use coding agents to work on research projects, I am trying to understand which decisions I can delegate: implementing a method, choosing the next experiment, interpreting a result, or reconsidering the research question itself.
One workflow I have used combines Codex on a server with ChatGPT as a second perspective on progress and next steps. I bring updates from the working environment into the discussion, then carry the resulting instructions back.
This gives me a place to examine decisions I am not yet sure how to make. It also leaves me responsible for connecting the two assistants. At times, I felt like their message bus.
I asked whether this counted as multi-agent collaboration, and why I should not let the server-side agent make its own decisions. Eventually, I put the wish plainly:
“I actually kind of hope it can run completely autonomously.”
That wish raises a practical question: how can I delegate more judgment while retaining a clear understanding of what the experiments establish? These are initial reflections from my own use, rather than a completed comparison of research-agent workflows.
Some of this friction may come from how the tools work together, and some from research questions I have not yet clarified. I cannot currently separate the two. These experiences describe what I encountered; they do not establish a general limitation of research agents.
The arrangement helped
I did not consult a second assistant only to supervise the first. Sometimes I did not know what to ask next. As I explained in our conversation, part of the reason I forwarded updates was that I did not know how to write a better prompt to continue.
Discussing the current state helped turn a vague concern into an actionable request. Should we reconstruct the project's history or continue experimenting? How much effort should go into the current blocker? What should the next step actually test?
Some judgment happened during the relay. Removing me would require the system to handle both message transfer and questions I had not yet expressed clearly.
I was still inside the workflow
Both assistants helped, but I still decided when to capture an update, which information to forward, and whether to pass the advice on unchanged.
The server-side agent had access to the working environment. The other assistant mainly saw what I sent it. I started wondering when this division provided another perspective and when it merely reinterpreted the same information.
I asked whether I needed to forward updates every time. If every stretch of progress required me to connect the two systems, increasingly automated code execution still tied my attention to the process. I wanted to reduce the ongoing effort of watching, understanding, relaying, and issuing instructions.
I wanted to let go, and worried about drift
In the same experience, I asked: “What exactly are we doing, and why does it feel as though we are drifting further off course?”
That was a feeling. This essay does not audit the complete code, experiments, and decision history, so I cannot present it as a verified instance of agent goal drift.
It nevertheless raised a question worth preserving: after an agent has done a great deal of work, how do I know it is still answering the question I care about?
Consider a hypothetical example. The question is whether failure detection improves a robot's closed-loop recovery. The agent trains a more accurate detector. That is useful progress. But without comparing recovery success, false alarms, and retry costs, the original question remains unanswered. This is not a description of my server experiments. It illustrates the judgment between finishing a local task and answering a research question.
Changing methods is not necessarily drift. If the agent discovers that the original approach fails and tries an alternative, it may be doing exactly what I want. I need to understand why it changed, what it is now testing, and how to inspect that change.
What I mean by autonomous
I do not want concern about mistakes to produce a workflow that asks me about every step. That would recreate the original problem.
To make the wish actionable, I want to try a division of responsibility: within a clear question and budget, let the agent diagnose the environment, revise implementation, choose small experiments, and continue from their results. When it believes the main research question or evaluation criterion should change, it should present a concrete reason and proposal.
I want it to be able to tell me that the original question may be mistaken, with supporting evidence and the cheapest next test.
Lvmin Zhang's Computation Shaped by Intention drew my attention to how intention enters a system and to the boundary between following and exploring. His exploring specifically means offering options awaiting acceptance. What I want to try also includes direct experimentation within authority granted beforehand. These meanings should remain distinct.
An interface worth trying would put the current question, available evidence, and reason for the next experiment together. I should not have to reconstruct the direction of research from a long stream of operational updates.
What a second agent could do
I still think another perspective may help. But I would rather test a second assistant that examines important conclusions than one that continuously supplies the next instruction.
Does this result support the claim? Is there a competing explanation? Has an intermediate metric become the final objective?
Two agents may agree because they read the same incomplete summary. My working view is that a second agent's value depends on what it actually checks. Its existence alone establishes little.
What I want to try next
My constraints are concrete: four 80 GB GPUs and approximately 32 hours per allocation. I want to start with a small, resumable task with a clear stopping condition and compare two workflows.
In the first, I continue relaying progress between assistants. In the second, I state the question, budget, and acceptance criteria upfront, let the server-side agent proceed, and receive inspectable updates at important decisions.
I want to record my actual time and whether the original question receives an answer, alongside completed work. A negative result can be an answer.
This comparison is not complete. If independent operation does not reduce my burden, or simply defers it to final review, I need to acknowledge that. If my intervention interrupts useful exploration, that deserves recording too.
I do not yet have a mature answer. The question comes from my experience, and it will continue to shape how I conduct research and use agents.
About me
I am Yaoliang Bian (Jeff), a doctoral student in the joint Westlake University–Shenzhen Loop Area Institute program. My interests include embodied intelligence, multimodal models, and research agents.
If you use agents to run experiments, I would like to hear concrete experiences: which judgments do you delegate, when do you intervene, and when has staying out of the way worked better?
让 AI 自主推进科研:我在尝试委托哪些判断?
用 AI 阅读与讨论 · Open in Claude / ChatGPT
把文章交给 AI,继续讨论其中的问题。若跳转后未带入提示词,请复制下方内容。
最近用 Coding Agent 推进科研,我开始思考哪些判断可以交给 AI:实现一个方法、选择下一项实验、解释结果,还是重新考虑研究问题本身?
我尝试过一种工作方式:让服务器上的 Codex 执行任务,再用 ChatGPT 讨论进展和下一步。我把工作环境里的信息带到讨论中,再把形成的指令传回去。
这种安排让我有机会梳理尚未想清楚的决策,也让我承担了连接两个助手的工作。有时,我感觉自己成了它们之间的 Prompt 中转站。
我问过这算不算多 Agent 协作,也问过为什么不让服务器上的 Codex 自己判断。最后,我说出了最直接的愿望:
“我其实有点希望他完全自主运行。”
这个愿望带来一个实际问题:怎样交出去更多研究判断,同时仍能清楚地理解实验说明了什么?以下是基于自己使用经历的初步思考,尚不是对不同科研 Agent 工作方式的完整实验比较。
这些摩擦可能同时来自工具的协作方式,以及研究问题本身尚未明确。现在我还无法区分两者,想通过后续实践继续观察。这些经历记录了我遇到的情况,还不足以说明科研 Agent 的普遍局限。
这套安排确实帮了我
我最初找第二个助手,不只是为了监督第一个。
有时我不知道下一步应该怎么问。我在对话里解释过:“其实我之前转给你也有因为我不知道继续下去怎么写出更好的 prompt。”
把当前状态拿出来讨论,能帮我把一个模糊的困惑变成具体要求:先查清历史状态,还是继续跑实验?眼前的阻塞值得解决多久?下一步到底要验证什么?
这说明所谓人工中转也不全是复制粘贴。有些判断就在这个过程中发生。如果直接把我删掉,系统需要接住的不只是消息传递,还有这些尚未表达清楚的问题。
但我没有真正退出工作流
两个助手都在帮我,我却仍然要决定什么时候截取进展、传递哪些信息,以及是否把建议原样交给服务器上的 Agent。
服务器上的 Agent 能访问工作环境,另一个助手收到的主要是我转来的内容。我开始怀疑:这样的分工在什么时候提供了第二种视角,又在什么时候只是把同一份信息重新解释一遍?
我也问过:“你觉得我应该每次都传递给你吗?”
如果每推进一段都要由我连接两边,那么即使代码执行越来越自动化,我的注意力仍然被绑在流程上。我真正想减少的是这种持续盯着、理解、转述、再发指令的负担。
我想放手,又担心走偏
同一段使用经历里,我还问过:
“我们到底在做什么为什么感觉越做越偏。”
这里需要讲清楚:这是我当时的感受。我没有在这篇文章里审计完整的代码、实验和决策记录,因此不能把它写成“Agent 已被证实发生目标漂移”。
但这种感觉很值得保留。它让我想问:怎样知道一个 Agent 做了很多工作以后,仍然在回答我真正关心的问题?
举一个假设例子。我的问题是“失败检测能否改善机器人的闭环恢复”。Agent 训练出了准确率更高的检测器,这是有用的进展。但如果我们始终没有比较恢复成功率、误报和重试成本,最初的问题仍没有答案。这个例子不是对我服务器实验的描述;它只是帮助我说明,局部任务完成与研究问题被回答之间,还隔着一步判断。
反过来,换方法不一定是走偏。一个 Agent 如果发现原方案不成立,主动尝试替代方案,可能正是我希望它具备的能力。真正需要辨别的是:它为什么改,修改后仍在检验什么,以及我能否看懂这次变化。
“完全自主”对我意味着什么
我不想因为担心出错,就把自主运行重新设计成每一步都要问我的流程。那会回到最初的问题。
为了把这个愿望变得可操作,我目前想尝试这样的分工:在问题和预算已经明确的范围内,让 Agent 自己排查环境、修改实现、选择小实验,并根据结果继续;当它认为应该更换主要研究问题或评价标准时,给我一份具体的理由和方案。
这不是要求它永远沿着最初计划走。相反,我希望它能提出“这个问题可能问错了”,但也说明依据是什么、接下来最便宜的验证是什么。
Lvmin Zhang 的 Computation Shaped by Intention 提醒我注意意图如何进入系统,以及 Following 和 Exploring 的边界。他文中的 Exploring 特指提供等待接受的选项;我这里想尝试的,还包括在提前授权的范围内直接做实验。这两种含义需要分开。
对我来说,一个值得尝试的界面,应该能把“当前要回答的问题”“已经有的证据”“下一步为什么值得做”放在一起。这样,我不必先翻完大量过程信息,才能判断研究在往哪里走。
第二个 Agent 应该做什么
我仍然觉得第二种视角可能有用。不过,比起让另一个助手不断给下一条指令,我更想尝试让它审查关键结论。
例如:这个结果有没有支持我们想说的话?是否遗漏了一个竞争解释?某个中间指标是不是被当成了最终目标?
两个 Agent 说法一致,也可能只是因为读了同一份不完整摘要。我的判断是,第二个 Agent 的价值需要来自它实际检查了什么,而不能仅仅来自“多了一个 Agent”。这仍是一个待验证的看法。
我准备怎样试
我的资源限制很具体:四张 80GB GPU,单次分配约 32 小时。因此,我想先选一个范围小、能恢复运行、也有明确结束条件的任务,比较两种工作方式。
第一种是继续由我在两个助手之间传递进展。第二种是一次交代问题、预算和验收条件,让服务器上的 Agent 自主推进,再在重要决策处给出可检查的更新。
我准备记录的,不只是做了多少事,还有我实际花了多少时间,以及最后有没有回答原来的问题。负结果也可以是答案。
这项比较还没做完。如果自主推进没有减少我的负担,或者只是把负担推迟到了最后检查,我就需要承认这次尝试没有达到目的。如果我介入反而打断了有效探索,也应该把它记录下来。
我还没有一个成熟的答案。这个问题来自我真实的使用经历,也会继续影响我怎样做研究、怎样使用 Agent。
关于我与交流
我是卞耀亮(Jeff),目前在西湖大学与深圳河套学院联合培养博士项目中,研究兴趣是具身智能、多模态模型和科研 Agent。
如果你也在让 Agent 跑实验,或者尝试过减少人类中转,我尤其想听到具体经历:你交给它哪些判断?什么情况下会介入?有没有一次“不插手”反而更好?