Return Label OCR: Which Large Language Model Comes Out on Top?

OCRMultimodalLLM EvaluationGWMSDeepSeek HarnessClaude Code

Why I ran this comparison

GWMS has a feature that lets warehouse staff photograph a return shipping label and extract the basic return information. The problem is that recognition takes far too long after the photo is taken.

So I obtained API keys from several providers and set up a comparison using the same prompt and the same API calling approach. I gave the task to both DeepSeek Harness and Claude Code CLI, asking each to run the evaluation and produce an HTML report.

Both received the same instruction, translated here:

Use the same prompt and the same API calling approach to test the models in ../OCRvsLLM. The .env file there contains API_KEYs for several large language models. Call them to perform recognition and produce a detailed HTML report. I want a side-by-side comparison of upload speed, recognition speed, and whether the recognition results agree. Also list the image dimensions and the size of the uploaded content.

Below are both reports exactly as delivered, with their authors identified. Their content and formatting are unchanged; the embedded originals are in Chinese.

Report one: produced by DeepSeek Harness

Report two: produced by Claude Code CLI

My assessment: the gap is clear

The same instruction produced two reports that were on entirely different levels.

First, they interpreted “recognition” differently. DeepSeek Harness took it to mean “describe what is in the images.” It invented a generic description prompt (“Please identify and describe the contents of the following four images…”) and ran it against conversational models such as gpt-4o and claude-sonnet-5. Claude Code CLI went back to the GWMS production code and found the prompt actually used by the live system in WmsReturnOrderServiceImpl#recognizeLabelFromImage. It used that production prompt, which asks for structured JSON fields, to compare the results field by field. Production needs fields such as customerCode, recipientRaw, and trackingNo, not a paragraph describing an image. These are fundamentally different tasks.

Second, the deliverables were different. I asked for “upload speed, recognition speed, whether the recognition results agree, image dimensions, and upload size.” DeepSeek Harness delivered one summary table: six models, timings, token counts, and a short text summary. Image dimensions and upload size were reduced to a single statement: “3.83 MB in total.” It did not compare whether the recognition results agreed at all. Claude Code CLI tested eight models across 108 calls. It scored each field in each image against a manually verified reference, with a maximum score of 24, rather than relying on a majority vote. It listed the pixel dimensions (608 × 1080) and three separate sizes: the original files, the Base64 data, and the actual request bodies. It also included a dedicated section comparing agreement between models, and even ran an additional experiment on converting images to JPEG and reducing their dimensions.

Third, the conclusions could even contradict each other. In the DeepSeek Harness report, deepseek-v4-flash-vision-exp appeared to work normally, with a recognition time of 14.21 seconds. In the Claude Code CLI report, it failed completely on two upside-down images, returning only its reasoning and no JSON. The difference comes down to what was being measured: a fluent image description does not mean the structured fields were extracted correctly. And photos taken with handheld warehouse terminals are often upside down.

Fourth, could the results be put to work directly? Claude Code CLI offered production recommendations ordered by how much benefit they could deliver with how little change: validate customerCode with a regular expression, check tracking-number check digits, convert images to JPEG on the client without reducing their dimensions, add a fallback for upside-down images, switch to structured output, and so on. DeepSeek Harness laid out its results but offered no conclusions that could guide improvements to the system.

Fifth, Claude Code CLI even uncovered the root cause of the slowness. After finishing the comparison, it went back and decompiled gwms-common-ai:1.0.20. It discovered that the enum name ModelType.QWEN_VL_MAX was misleading: the model string actually sent to Alibaba Cloud was qwen3.7-plus, not qwen-vl-max. In other words, the model we had been treating as the “production baseline” was not the one running in production at all. The actual production model, qwen3.7-plus, took 8.07 seconds for inference and 11.78 seconds end to end. Switching to qwen-vl-max brought inference down to 2.88 seconds, a 64% reduction, with the same 94% accuracy. That was the main contributor to the slow recognition response, and the DeepSeek Harness report had not come close to finding it.

The gap is clear. Given the same instruction, one delivered a readable summary table; the other delivered an evaluation that could directly guide production changes, and uncovered a production configuration problem along the way. This is about more than how capable the model is. It is also about whether, upon receiving the task, the agent first works out what this “recognition” feature actually needs to produce in production, and what happens when it gets things wrong.

← Back to blog