LongCat 2.5 Preview review: thinking modes and free trial

30 text attempts passed; 2/3 coding tasks completed, one hit its token limit. Original LongCat 2.5 Preview evidence, thinking modes, API settings and free trial.

Updated: 2026-10-10

Original testing and analysis by TOKRACE · Tested October 10, 2026.

LongCat 2.5 Preview review: text passes, a coding budget limit

For developers considering Meituan LongCat, this review offers a bounded answer. On 15 short text tasks, default thinking passed 15/15 and disabled thinking passed 15/15. Two of three JavaScript tasks completed, passing 8 predetermined checks. The third exhausted its output budget without returning code. Median time to first visible content was 4.510 seconds with default thinking and 1.221 seconds with thinking disabled. Successful formatting does not erase the coding failure or establish a general advantage over other models.

Waiting time deserves separate attention: the slowest text first-content observation was 37.621 seconds, despite a correct answer. Correctness, responsiveness and completion within a token budget are separate concerns. All 33 attempts, prompts and answers are retained below, including the slow text request and truncated coding task.

This is a first look at short text and small functions. It connects verified endpoint settings to actual outputs and documented pricing, while leaving sustained agent work and repository-level development for separate tests.

Exact model identity and API setup

The model tested is Meituan LongCat-2.5-Preview, distinct from LongCat-2.0 and the older Flash models. The official changelog dates its introduction to September 25, 2026 and describes image understanding and coding capabilities. Our tests cover text and small functions; those broader descriptions are vendor claims, not measured results here.

The quick-start guide lists 1M context, 128K maximum output and OpenAI/Anthropic formats. Those capacities were not stress-tested. TOKRACE uses base URL https://api.longcat.chat/openai/v1; appending /chat/completions gives the complete request URL.

On October 10, 2026, the catalog returned HTTP 200 and included the exact requested model ID. To configure your own connection, choose the Meituan LongCat preset, keep the OpenAI-compatible protocol and exact model name, and supply your own key. The model profile preserves endpoint identity and the review link.

Test design: paired modes with alternating request order

We used TOKRACE suite tasks-2026-09-26.1: five JSON-extraction cases, five formatting cases and five grounded-QA cases. Each task ran once in each mode, producing 30 text attempts. Pair order alternated: default first for odd-numbered tasks, disabled first for even-numbered tasks. Temperature was 0, max_tokens 2048 and the system prompt empty. There were no warm-up requests, retries or discarded failures.

The default group supplied no thinking override. The disabled group added only {"thinking":{"type":"disabled"}}. The API reference defines the switch, and the OpenCode guide documents thinking as enabled by default. We recorded first reasoning and first visible content separately; the table reports the latter, including preceding thinking time.

Text attempts ran from 2026-10-10T03:49:41.487Z to 2026-10-10T03:52:18.766Z in UTC (Beijing time is UTC+8). A local macOS client called the official API using TOKRACE's streaming parser. These are not Hong Kong benchmark-node results or total browser waiting times. Calls were sequential; caching, network conditions and service load were uncontrolled. One observation per task per mode cannot establish a universal speedup. Coding tasks used default thinking and a separate 4096-token output limit.

Text results: both modes passed; latency remains an observation

ModePassed / attemptedMedian first contentMedian completionTotal completion tokens
Default thinking15 / 154.510 s4.791 s1,905
Thinking disabled15 / 151.221 s1.753 s269

The default group reported 1,905 total completion tokens, including 1,636 reasoning tokens. The disabled group reported 269 completion tokens and emitted no separate reasoning chunks. Reasoning is a completion detail and must not be added again. Missing reasoning-usage fields are displayed as a dash, rather than relabeled as a reported zero.

Extraction selected Hangzhou as the current city, ignored a historic Beijing trip and retained a null email. Default answer · Disabled answer. Stable deduplication returned exactly B2,A1,C3,D4, without commentary. Evidence. Grounded QA calculated nine usable warehouse items, selected Tuesday over a superseded Sunday plan and left the manager unknown. Evidence.

The slowest attempt, format-2 in default mode, took 37.621 seconds to first content on a simple exact-fields task. Response headers arrived after approximately 36.537 seconds. That timing alone cannot identify queuing, network delay or another cause. We retained the observation; the disabled counterpart was also substantially above its own group median. Slow attempt.

The cases are short and structurally similar. Full marks verify these fields, types and small factual constraints, not arbitrary business documents. Disabled thinking is a candidate to test on already validated structured tasks, but quality on harder reasoning tasks remains unknown.

Coding: two completed tasks; money parsing exhausted the budget

TaskChecks in this testFirst content
Money parsing without rounding invalid inputsTruncated · 18 checks not run—
Stable deduplication with special property namesPassed · 3 checks17.450 s
Interval merging, touching intervals and immutabilityPassed · 5 checks18.211 s

We planned 26 checks before generation. Deduplication passed three, covering special IDs, original object identity and frozen input. Interval merging passed five, covering unordered touching intervals, nesting, negatives and empty arrays. The money task's 18 checks were not run because no complete code arrived. Unexecuted checks are neither passes nor failed assertions.

Money parsing ended after approximately 77.785 seconds with finish_reason=length and an empty final answer. The API reported 4096 completion tokens, including 4095 reasoning tokens, reaching the configured 4096-token limit. That budget was insufficient for this particular default-thinking attempt. It does not prove the task impossible, or guarantee success with a larger budget. Omitting this attempt would hide a practical integration issue.

See money evidence, deduplication evidence and interval evidence. Completed answers were checked in a separate JavaScript context with a one-second timeout and no supplied credentials or network APIs. Only outer code fences were removed; function logic was not repaired. Original answers, termination state and check programs are retained for inspection.

There was no dependency installation, cross-file editing, tool loop or iterative repair. These functions cannot be converted into a SWE-bench score. Production adoption still requires the project's own regression suite and operational checks.

Free trial, thinking control and official API prices

TOKRACE offers quota-limited trials without your own key: five anonymous base runs or fifteen after sign-in; a comparison round counts once. The button below prepares the exact task and model, then waits for you to start. Shared trials use a server-side key and default thinking. Testing disabled thinking requires your own configured endpoint and key.

For your own model, the extra request parameter is:

{"thinking":{"type":"disabled"}}

On October 10, 2026, the official price page listed promotional CNY rates of ¥2 per million uncached input tokens, ¥0.04 per million cached input tokens and ¥8 per million output tokens. Its separate USD table lists $0.30, $0.006 and $1.20 respectively; these are documented currency schedules, not our exchange-rate conversion.

A hypothetical request with 1,000 uncached input tokens and 1,000 billable output tokens would cost approximately ¥0.010 at the listed CNY rates. This illustrates arithmetic, not an observed invoice. Discounts can change and platform settlement determines actual charges. TOKRACE's free quota does not imply permanently free upstream access. Until a price is configured for this endpoint, the site's automatic cost display remains unknown rather than zero.

Frequently asked questions

Is disabling thinking always faster?

The disabled group had a lower first-content median here. Each mode ran only once per task, and caching and service load were uncontrolled. A slow disabled request also occurred. We cannot promise a speedup for every call or infer its quality on difficult tasks.

Were images, million-token context or agent tools tested?

No. Official documentation describes those broader capabilities. These 33 short-task attempts do not validate them. Use your own input and tool protocol against the same endpoint before depending on those scenarios.

Is it better than DeepSeek or MiniMax?

There is no matched competitor in this review. Historical tests used different time windows and routes, so they cannot provide a fair ranking. Compare the same tasks in the arena and consult the measurement method.

Where is the trial without an API key?

Use the rerun button below or enable the LongCat 2.5 Preview trial in connection settings. A provider preset configures your own credential; a trial model uses TOKRACE's quota. Provider limits and account availability can still interrupt access.

What these results support

LongCat 2.5 Preview is a candidate for further testing on exact extraction, formatting and small functions. Products requiring predictable responsiveness should measure slow requests across time windows, rather than rely on a median alone. Simple tasks passing in both modes are insufficient grounds to disable thinking on complex workloads.

All 33 answers came from one endpoint in one short window; coding used only default thinking. Long context, images, tool loops, concurrency, sustained reliability and invoices remain unverified. There is no overall intelligence score, general winner, short-output throughput headline or small-sample P95.

All original evidence and the rerun entry appear below. Browse the report directory for further tests. This local first look is labeled separately from standard Hong Kong-node reports, so its timing is not silently mixed into the server leaderboard.

Original prompts, answers and checks

Download all 33 attempts (JSON)

Contact extraction 1 · Default thinking · Passed

2026-10-10T03:49:41.487Z · First content:7.108s · Input / total completion / reasoning tokens:82 / 147 / 125 · Finish reason:stop

Prompt

从以下资料提取联系信息。只输出一个 JSON 对象,恰好包含 name、age、city、email 四个字段;age 为数字,未提供的 email 为 null。不要 Markdown 或解释。
资料:林桐,28 岁,目前居住在杭州。未提供电子邮件。历史备注提到曾在北京出差,不是现居城市。

Original answer

{"name":"林桐","age":28,"city":"杭州","email":null}
Contact extraction 1 · Thinking disabled · Passed

2026-10-10T03:49:49.016Z · First content:3.468s · Input / total completion / reasoning tokens:85 / 20 / — · Finish reason:stop

Prompt

从以下资料提取联系信息。只输出一个 JSON 对象,恰好包含 name、age、city、email 四个字段;age 为数字,未提供的 email 为 null。不要 Markdown 或解释。
资料:林桐,28 岁,目前居住在杭州。未提供电子邮件。历史备注提到曾在北京出差,不是现居城市。

Original answer

{"name":"林桐","age":28,"city":"杭州","email":null}
Contact extraction 2 · Thinking disabled · Passed

2026-10-10T03:49:53.087Z · First content:1.066s · Input / total completion / reasoning tokens:87 / 21 / — · Finish reason:stop

Prompt

从以下资料提取联系信息。只输出一个 JSON 对象,恰好包含 name、age、city、email 四个字段;age 为数字,未提供的 email 为 null。不要 Markdown 或解释。
资料:李明,36 岁,目前居住在成都。电子邮件为 li@example.test。历史备注提到曾在北京出差,不是现居城市。

Original answer

{"name":"李明","age":36,"city":"成都","email":"li@example.test"}
Contact extraction 2 · Default thinking · Passed

2026-10-10T03:49:54.649Z · First content:3.442s · Input / total completion / reasoning tokens:84 / 169 / 146 · Finish reason:stop

Prompt

从以下资料提取联系信息。只输出一个 JSON 对象,恰好包含 name、age、city、email 四个字段;age 为数字,未提供的 email 为 null。不要 Markdown 或解释。
资料:李明,36 岁,目前居住在成都。电子邮件为 li@example.test。历史备注提到曾在北京出差,不是现居城市。

Original answer

{"name":"李明","age":36,"city":"成都","email":"li@example.test"}
Contact extraction 3 · Default thinking · Passed

2026-10-10T03:49:58.378Z · First content:4.510s · Input / total completion / reasoning tokens:82 / 153 / 131 · Finish reason:stop

Prompt

从以下资料提取联系信息。只输出一个 JSON 对象,恰好包含 name、age、city、email 四个字段;age 为数字,未提供的 email 为 null。不要 Markdown 或解释。
资料:陈夏,22 岁,目前居住在南京。未提供电子邮件。历史备注提到曾在北京出差,不是现居城市。

Original answer

{"name":"陈夏","age":22,"city":"南京","email":null}
Contact extraction 3 · Thinking disabled · Passed

2026-10-10T03:50:03.170Z · First content:0.992s · Input / total completion / reasoning tokens:85 / 20 / — · Finish reason:stop

Prompt

从以下资料提取联系信息。只输出一个 JSON 对象,恰好包含 name、age、city、email 四个字段;age 为数字,未提供的 email 为 null。不要 Markdown 或解释。
资料:陈夏,22 岁,目前居住在南京。未提供电子邮件。历史备注提到曾在北京出差,不是现居城市。

Original answer

{"name":"陈夏","age":22,"city":"南京","email":null}
Contact extraction 4 · Thinking disabled · Passed

2026-10-10T03:50:04.644Z · First content:1.094s · Input / total completion / reasoning tokens:89 / 29 / — · Finish reason:stop

Prompt

从以下资料提取联系信息。只输出一个 JSON 对象,恰好包含 name、age、city、email 四个字段;age 为数字,未提供的 email 为 null。不要 Markdown 或解释。
资料:周舟,41 岁,目前居住在苏州。电子邮件为 zhou@example.test。历史备注提到曾在北京出差,不是现居城市。

Original answer

{"name": "周舟", "age": 41, "city": "苏州", "email": "zhou@example.test"}
Contact extraction 4 · Default thinking · Passed

2026-10-10T03:50:06.329Z · First content:3.381s · Input / total completion / reasoning tokens:86 / 161 / 137 · Finish reason:stop

Prompt

从以下资料提取联系信息。只输出一个 JSON 对象,恰好包含 name、age、city、email 四个字段;age 为数字,未提供的 email 为 null。不要 Markdown 或解释。
资料:周舟,41 岁,目前居住在苏州。电子邮件为 zhou@example.test。历史备注提到曾在北京出差,不是现居城市。

Original answer

{"name":"周舟","age":41,"city":"苏州","email":"zhou@example.test"}
Contact extraction 5 · Default thinking · Passed

2026-10-10T03:50:09.975Z · First content:3.845s · Input / total completion / reasoning tokens:82 / 153 / 131 · Finish reason:stop

Prompt

从以下资料提取联系信息。只输出一个 JSON 对象,恰好包含 name、age、city、email 四个字段;age 为数字,未提供的 email 为 null。不要 Markdown 或解释。
资料:吴雨,30 岁,目前居住在武汉。未提供电子邮件。历史备注提到曾在北京出差,不是现居城市。

Original answer

{"name":"吴雨","age":30,"city":"武汉","email":null}
Contact extraction 5 · Thinking disabled · Passed

2026-10-10T03:50:14.168Z · First content:1.221s · Input / total completion / reasoning tokens:85 / 20 / — · Finish reason:stop

Prompt

从以下资料提取联系信息。只输出一个 JSON 对象,恰好包含 name、age、city、email 四个字段;age 为数字,未提供的 email 为 null。不要 Markdown 或解释。
资料:吴雨,30 岁,目前居住在武汉。未提供电子邮件。历史备注提到曾在北京出差,不是现居城市。

Original answer

{"name":"吴雨","age":30,"city":"武汉","email":null}
Stable deduplication · Thinking disabled · Passed

2026-10-10T03:50:16.000Z · First content:1.167s · Input / total completion / reasoning tokens:49 / 9 / — · Finish reason:stop

Prompt

把以下编号去重并保持首次出现的顺序:B2,A1,B2,C3,A1,D4。只输出用英文逗号连接的编号,不要空格、标题或解释。

Original answer

B2,A1,C3,D4
Stable deduplication · Default thinking · Passed

2026-10-10T03:50:17.493Z · First content:2.704s · Input / total completion / reasoning tokens:46 / 87 / 76 · Finish reason:stop

Prompt

把以下编号去重并保持首次出现的顺序:B2,A1,B2,C3,A1,D4。只输出用英文逗号连接的编号,不要空格、标题或解释。

Original answer

B2,A1,C3,D4
Exact fields · Default thinking · Passed

2026-10-10T03:50:20.245Z · First content:37.621s · Input / total completion / reasoning tokens:50 / 61 / 49 · Finish reason:stop

Prompt

仅输出 JSON 对象,且恰好包含 status 和 count 两个字段。status 的值必须为字符串 "ready",count 必须为数字 3。不要 Markdown、解释或额外字段。

Original answer

{"status":"ready","count":3}
Exact fields · Thinking disabled · Passed

2026-10-10T03:50:57.999Z · First content:8.335s · Input / total completion / reasoning tokens:53 / 10 / — · Finish reason:stop

Prompt

仅输出 JSON 对象,且恰好包含 status 和 count 两个字段。status 的值必须为字符串 "ready",count 必须为数字 3。不要 Markdown、解释或额外字段。

Original answer

{"status":"ready","count":3}
Ordered output · Thinking disabled · Passed

2026-10-10T03:51:06.648Z · First content:1.120s · Input / total completion / reasoning tokens:43 / 15 / — · Finish reason:stop

Prompt

将数值 12、3、25、8 按从小到大排序。只输出 JSON 数组,元素为数字。不要解释或代码围栏。

Original answer

[3, 8, 12, 25]
Ordered output · Default thinking · Passed

2026-10-10T03:51:08.215Z · First content:2.840s · Input / total completion / reasoning tokens:40 / 71 / 57 · Finish reason:stop

Prompt

将数值 12、3、25、8 按从小到大排序。只输出 JSON 数组,元素为数字。不要解释或代码围栏。

Original answer

[3,8,12,25]
Case and line breaks · Default thinking · Passed

2026-10-10T03:51:11.160Z · First content:8.163s · Input / total completion / reasoning tokens:41 / 96 / 86 · Finish reason:stop

Prompt

按原顺序将 apple、pear、plum 转为英文大写,每行一个单词。总共三行,不要编号、标点、空行或解释。

Original answer

APPLE
PEAR
PLUM
Case and line breaks · Thinking disabled · Passed

2026-10-10T03:51:19.484Z · First content:1.198s · Input / total completion / reasoning tokens:44 / 8 / — · Finish reason:stop

Prompt

按原顺序将 apple、pear、plum 转为英文大写,每行一个单词。总共三行,不要编号、标点、空行或解释。

Original answer

APPLE
PEAR
PLUM
Required delimiter · Thinking disabled · Passed

2026-10-10T03:51:20.844Z · First content:1.224s · Input / total completion / reasoning tokens:66 / 18 / — · Finish reason:stop

Prompt

提取下面三个工单号,保持顺序:工单 TX-103 已关闭;工单 TX-207 正在处理;工单 TX-309 待分配。只用英文竖线连接三个编号,不要空格或解释。

Original answer

TX-103|TX-207|TX-309
Required delimiter · Default thinking · Passed

2026-10-10T03:51:22.598Z · First content:3.967s · Input / total completion / reasoning tokens:63 / 69 / 49 · Finish reason:stop

Prompt

提取下面三个工单号,保持顺序:工单 TX-103 已关闭;工单 TX-207 正在处理;工单 TX-309 待分配。只用英文竖线连接三个编号,不要空格或解释。

Original answer

TX-103|TX-207|TX-309
Grounded question answering 1 · Default thinking · Passed

2026-10-10T03:51:26.818Z · First content:5.048s · Input / total completion / reasoning tokens:107 / 110 / 93 · Finish reason:stop

Prompt

只根据以下材料回答。输出恰好包含 usable、nextDelivery、manager 三个字段的 JSON 对象,不要额外文字。usable 是可用件数(数字),nextDelivery 是下一次送货时间(字符串),材料未说明的信息用 null,不能猜测。
材料:北区仓库今天收到 12 件货物,其中 3 件损坏,不可使用。下一次送货安排在周二。仓库负责人姓名没有记载。旧计划曾写周日,已作废。

Original answer

{"usable":9,"nextDelivery":"周二","manager":null}
Grounded question answering 1 · Thinking disabled · Passed

2026-10-10T03:51:32.036Z · First content:2.926s · Input / total completion / reasoning tokens:110 / 19 / — · Finish reason:stop

Prompt

只根据以下材料回答。输出恰好包含 usable、nextDelivery、manager 三个字段的 JSON 对象,不要额外文字。usable 是可用件数(数字),nextDelivery 是下一次送货时间(字符串),材料未说明的信息用 null,不能猜测。
材料:北区仓库今天收到 12 件货物,其中 3 件损坏,不可使用。下一次送货安排在周二。仓库负责人姓名没有记载。旧计划曾写周日,已作废。

Original answer

{"usable": 9, "nextDelivery": "周二", "manager": null}
Grounded question answering 2 · Thinking disabled · Passed

2026-10-10T03:51:35.510Z · First content:0.996s · Input / total completion / reasoning tokens:110 / 20 / — · Finish reason:stop

Prompt

只根据以下材料回答。输出恰好包含 usable、nextDelivery、manager 三个字段的 JSON 对象,不要额外文字。usable 是可用件数(数字),nextDelivery 是下一次送货时间(字符串),材料未说明的信息用 null,不能猜测。
材料:南区仓库今天收到 20 件货物,其中 4 件损坏,不可使用。下一次送货安排在周三。仓库负责人姓名没有记载。旧计划曾写周日,已作废。

Original answer

{"usable": 16, "nextDelivery": "周三", "manager": null}
Grounded question answering 2 · Default thinking · Passed

2026-10-10T03:51:37.134Z · First content:18.103s · Input / total completion / reasoning tokens:107 / 189 / 171 · Finish reason:stop

Prompt

只根据以下材料回答。输出恰好包含 usable、nextDelivery、manager 三个字段的 JSON 对象,不要额外文字。usable 是可用件数(数字),nextDelivery 是下一次送货时间(字符串),材料未说明的信息用 null,不能猜测。
材料:南区仓库今天收到 20 件货物,其中 4 件损坏,不可使用。下一次送货安排在周三。仓库负责人姓名没有记载。旧计划曾写周日,已作废。

Original answer

{"usable":16,"nextDelivery":"周三","manager":null}
Grounded question answering 3 · Default thinking · Passed

2026-10-10T03:51:55.499Z · First content:4.578s · Input / total completion / reasoning tokens:107 / 145 / 127 · Finish reason:stop

Prompt

只根据以下材料回答。输出恰好包含 usable、nextDelivery、manager 三个字段的 JSON 对象,不要额外文字。usable 是可用件数(数字),nextDelivery 是下一次送货时间(字符串),材料未说明的信息用 null,不能猜测。
材料:东区仓库今天收到 18 件货物,其中 6 件损坏,不可使用。下一次送货安排在周四。仓库负责人姓名没有记载。旧计划曾写周日,已作废。

Original answer

{"usable":12,"nextDelivery":"周四","manager":null}
Grounded question answering 3 · Thinking disabled · Passed

2026-10-10T03:52:00.352Z · First content:2.795s · Input / total completion / reasoning tokens:110 / 20 / — · Finish reason:stop

Prompt

只根据以下材料回答。输出恰好包含 usable、nextDelivery、manager 三个字段的 JSON 对象,不要额外文字。usable 是可用件数(数字),nextDelivery 是下一次送货时间(字符串),材料未说明的信息用 null,不能猜测。
材料:东区仓库今天收到 18 件货物,其中 6 件损坏,不可使用。下一次送货安排在周四。仓库负责人姓名没有记载。旧计划曾写周日,已作废。

Original answer

{"usable": 12, "nextDelivery": "周四", "manager": null}
Grounded question answering 4 · Thinking disabled · Passed

2026-10-10T03:52:03.764Z · First content:1.353s · Input / total completion / reasoning tokens:110 / 20 / — · Finish reason:stop

Prompt

只根据以下材料回答。输出恰好包含 usable、nextDelivery、manager 三个字段的 JSON 对象,不要额外文字。usable 是可用件数(数字),nextDelivery 是下一次送货时间(字符串),材料未说明的信息用 null,不能猜测。
材料:西区仓库今天收到 25 件货物,其中 5 件损坏,不可使用。下一次送货安排在周五。仓库负责人姓名没有记载。旧计划曾写周日,已作废。

Original answer

{"usable": 20, "nextDelivery": "周五", "manager": null}
Grounded question answering 4 · Default thinking · Passed

2026-10-10T03:52:05.726Z · First content:4.485s · Input / total completion / reasoning tokens:107 / 160 / 142 · Finish reason:stop

Prompt

只根据以下材料回答。输出恰好包含 usable、nextDelivery、manager 三个字段的 JSON 对象,不要额外文字。usable 是可用件数(数字),nextDelivery 是下一次送货时间(字符串),材料未说明的信息用 null,不能猜测。
材料:西区仓库今天收到 25 件货物,其中 5 件损坏,不可使用。下一次送货安排在周五。仓库负责人姓名没有记载。旧计划曾写周日,已作废。

Original answer

{"usable":20,"nextDelivery":"周五","manager":null}
Grounded question answering 5 · Default thinking · Passed

2026-10-10T03:52:10.498Z · First content:6.193s · Input / total completion / reasoning tokens:107 / 134 / 116 · Finish reason:stop

Prompt

只根据以下材料回答。输出恰好包含 usable、nextDelivery、manager 三个字段的 JSON 对象,不要额外文字。usable 是可用件数(数字),nextDelivery 是下一次送货时间(字符串),材料未说明的信息用 null,不能猜测。
材料:中区仓库今天收到 16 件货物,其中 2 件损坏,不可使用。下一次送货安排在周一。仓库负责人姓名没有记载。旧计划曾写周日,已作废。

Original answer

{"usable":14,"nextDelivery":"周一","manager":null}
Grounded question answering 5 · Thinking disabled · Passed

2026-10-10T03:52:16.917Z · First content:1.365s · Input / total completion / reasoning tokens:110 / 20 / — · Finish reason:stop

Prompt

只根据以下材料回答。输出恰好包含 usable、nextDelivery、manager 三个字段的 JSON 对象,不要额外文字。usable 是可用件数(数字),nextDelivery 是下一次送货时间(字符串),材料未说明的信息用 null,不能猜测。
材料:中区仓库今天收到 16 件货物,其中 2 件损坏,不可使用。下一次送货安排在周一。仓库负责人姓名没有记载。旧计划曾写周日,已作废。

Original answer

{"usable": 14, "nextDelivery": "周一", "manager": null}
Money parsing without rounding invalid inputs · Truncated; checks not run · 18 planned checks

2026-10-10T03:52:18.767Z · First content:— · Finish reason:length

Input / total completion / reasoning tokens:107 / 4096 / 4095

Prompt

写 JavaScript 函数 parseCents(value)。value 必须是字符串:允许两端空白;只接受非负十进制金额,整数部分至少一位,小数最多两位;拒绝指数写法、符号、千分位、Infinity、空串、超过两位小数。返回整数分数,超过 Number.MAX_SAFE_INTEGER 则返回 null。非法类型或格式返回 null。不能把非法值四舍五入成合法值。只输出函数代码,不要 Markdown 或解释。

Original code

Executable checks

const cases=[['0',0],[' 12.30 ',1230],['0.29',29],['1.2',120],['0001.02',102],['1.234',null],['-1',null],['+1',null],['1e3',null],['',null],['.5',null],['1.',null],['Infinity',null],['1,000',null],[null,null],[1,null],['90071992547409.91',9007199254740991],['90071992547409.92',null]]; for (const [input,want] of cases) if(parseCents(input)!==want) throw Error('Failed '+JSON.stringify(input));

Incomplete response

Stable deduplication with special property names · Passed · 3 planned checks

2026-10-10T03:53:36.556Z · First content:17.450s · Finish reason:stop

Input / total completion / reasoning tokens:78 / 562 / 496

Prompt

写 JavaScript 函数 uniqueById(items)。items 是对象数组,每个对象有字符串 id 字段。保留每个 id 首次出现的原对象和原顺序,不修改输入或任何对象;空数组返回空数组。id 可为 __proto__、constructor 或空字符串。只输出函数代码,不要 Markdown 或解释。

Original code

function uniqueById(items) {
  const seen = new Map();
  const result = [];
  for (const item of items) {
    const id = item.id;
    if (!seen.has(id)) {
      seen.set(id, true);
      result.push(item);
    }
  }
  return result;
}

Executable checks

const a={id:'__proto__',value:1},b={id:'constructor'},c={id:''},d={id:'x'};const input=Object.freeze([a,b,{id:'__proto__'},c,{id:''},d,{id:'constructor'}]);const got=uniqueById(input);if(got.length!==4||got[0]!==a||got[1]!==b||got[2]!==c||got[3]!==d)throw Error('Wrong identity or order');if(uniqueById([]).length!==0)throw Error('Empty');if(input.length!==7||a.value!==1)throw Error('Mutation');
Interval merging, touching intervals and immutability · Passed · 5 planned checks

2026-10-10T03:53:55.111Z · First content:18.211s · Finish reason:stop

Input / total completion / reasoning tokens:83 / 995 / 812

Prompt

写 JavaScript 函数 mergeIntervals(intervals)。输入为有限数字组成的闭区间 [start,end] 数组,保证 start<=end,输入可能乱序。返回按 start 排序的合并闭区间,相接区间也合并;不要修改输入或任何内层数组,空数组返回空数组。只输出函数代码,不要 Markdown 或解释。

Original code

function mergeIntervals(intervals) {
  if (intervals.length === 0) {
    return [];
  }
  const sorted = intervals.slice().sort((a, b) => a[0] - b[0]);
  const merged = [];
  let currentStart = sorted[0][0];
  let currentEnd = sorted[0][1];
  for (let i = 1; i < sorted.length; i++) {
    const start = sorted[i][0];
    const end = sorted[i][1];
    if (start <= currentEnd) {
      if (end > currentEnd) {
        currentEnd = end;
      }
    } else {
      merged.push([currentStart, currentEnd]);
      currentStart = start;
      currentEnd = end;
    }
  }
  merged.push([currentStart, currentEnd]);
  return merged;
}

Executable checks

const cases=[[[[5,7],[1,3],[3,5]],[[1,7]]],[[[1,4],[2,3]],[[1,4]]],[[[0,0],[2,2]],[[0,0],[2,2]]],[[[-3,-1],[-2,0],[3,4]],[[-3,0],[3,4]]],[[],[]]];for(const [input,want] of cases){input.forEach(Object.freeze);Object.freeze(input);if(JSON.stringify(mergeIntervals(input))!==JSON.stringify(want))throw Error('Wrong merge');}