Oct 9 edition/Reporting & analysis
AgentsResearchBusiness

AgentsAutonomy & tool use

NVIDIA's KGMON team explains the agent harness behind its second place in the KDD Cup 2026 Data Agents track

NVIDIA's KGMON team placed second in the KDD Cup 2026 Data Agents track with a mandated Qwen3.5-35B-A3B model. It improved its result by engineering the surrounding harness: unified SQL access, a small toolset, schema preflight checks and full run logging.

THE CORE IDEAS4 TAKEAWAYS
01

With the model fixed by the rules, KGMON improved the code around it. It loaded CSV and JSON data into one SQLite database per task and limited the agent to a few tools. A preflight step mapped join keys, units and missing values before each task, and prose and video were pre-extracted outside the agent loop. [1]

02

On DataSpace, the competition's 410-task benchmark, changing only the harness moved accuracy by 15.36 points with the same model. The best of six frontier models reached 66.34%, and cross-format evidence and joins lowered accuracy for every model. These are the two weak points KGMON's design targeted. [2] [1]

03

Automated self-improvement loops overfit the benchmark. They hardcoded training examples into prompts, produced contradictory instructions and added brittle post-processing rules. The team kept a change only after it passed held-out tasks, a prompt audit and human approval. [1]

04

The first-place team's code is public with Docker support. Its repository reports a self-reported mean score of 0.8033 over 10 test runs, and slides and recordings from eight featured teams are also available. [4] [5]

WHY IT MATTERS

the contest benchmark found that harness changes alone shifted accuracy by 15.36 points with one model. Implication: teams running small open models may gain more from cleaner data access, fewer tools and traceable runs than from upgrading models.

Executive brief

NVIDIA's KGMON team took second place in the KDD Cup 2026 Data Agents competition without being allowed to choose its model. The rules fixed the LLM (Qwen3.5-35B-A3B), so the team won its placing by improving the "harness," the code around the model. It merged all structured data into one SQLite database, gave the agent only a few tools, checked the schema before each task and logged every run for review (NVIDIA). The benchmark behind the contest points the same way: with the model held constant, changing the harness moved accuracy by 15.36 points (DataSpace, arXiv). NVIDIA has not published KGMON's own score.

What changed and event timeline

  1. Phase-1 demo data released

    The HKUST-organized Data Agents track opens with a $30,000 prize pool. It sits beside a much larger recommendation-systems track ().

  2. Official benchmark paper posted

    DataSpace, the competition's evaluation benchmark, appears on arXiv. It covers 410 tasks across CSV, JSON, SQLite, Markdown, PDF and video ().

  3. Also

    Date unconfirmed: Winning solution published

    The first-place team publishes code on GitHub. It reports a mean score of 0.8033 (standard deviation 0.0252) over 10 test runs ().

  4. NVIDIA publishes its playbook

    KGMON describes the design behind its second place and links slides and recordings from eight featured teams (;).

Capabilities and access

  • Model: Qwen3.5-35B-A3B, an open model the competition required (NVIDIA).
  • Tools the agent could call: schema(), sql(query), write_answer(df) and prose_helper() in two modes, answer and table.
  • Competition setting: no internet access, mixed task bundles, scoring on exact values, and a long overall run budget.
Read the full section
  • Model: Qwen3.5-35B-A3B, an open model the competition required (NVIDIA).
  • Tools the agent could call: schema(), sql(query), write_answer(df) and prose_helper() in two modes, answer and table.
  • Competition setting: no internet access, mixed task bundles, scoring on exact values, and a long overall run budget.
  • Code: NVIDIA has not released KGMON's code. The first-place code is public and ships with Docker (GitHub).

Technical analysis

  • CSV and JSON files were loaded as tables in the task's SQLite database.
  • A Python environment kept variables between tool calls, and middleware fixed malformed tool calls so one bad call did not end a run.
  • The agent could not read whole files.
Read the full section
  • One query surface. CSV and JSON files were loaded as tables in the task's SQLite database. A "preflight" step then mapped join keys, look-alike fields, units, missing-value patterns and what each row represents, and gave that to the agent before it started.
  • Persistent execution. A Python environment kept variables between tool calls, and middleware fixed malformed tool calls so one bad call did not end a run.
  • Prose kept separate. The agent could not read whole files. It used limited previews and regex search, then passed relevant chunks to a separate LLM call (temperature 0, reasoning off) that returned a short answer or a SQL table.
  • Video handled up front. Keyframes were extracted and lined up with the audio transcript before the agent loop began.
  • Choosing among runs. Repeated attempts were grouped by answer values rather than column names.
  • Scoring. DataSpace matches columns regardless of header names, normalizes types and precision, and compares rows in order (arXiv).

Claims and evidence

  • DataSpace reports a 15.36-point accuracy spread across harnesses with the same model, and a best accuracy of 66.34% across six frontier models (arXiv).
  • KGMON's techniques reduced errors. NVIDIA says unified SQL cut routing failures and schema scouting cut wrong-column errors (NVIDIA).
  • First-place score. The 0.8033 mean is self-reported in the winner's repository (GitHub). No independent leaderboard source was found.
Read the full section
  • Harness design matters a lot. DataSpace reports a 15.36-point accuracy spread across harnesses with the same model, and a best accuracy of 66.34% across six frontier models (arXiv). These are the organizers' own numbers, not the contest leaderboard.
  • KGMON's techniques reduced errors. NVIDIA says unified SQL cut routing failures and schema scouting cut wrong-column errors (NVIDIA). It gives no ablations or numbers, so each technique's effect cannot be measured from public material.
  • First-place score. The 0.8033 mean is self-reported in the winner's repository (GitHub). No independent leaderboard source was found.
  • No independent check of KGMON's placing. Other coverage (HyperAI) restates NVIDIA's post.

Context and prior work

DataSpace builds on a line of data-science agent research such as AutoMind. Its distinguishing feature is mixed formats, including video. Its own results show that combining evidence across formats and performing joins lowered accuracy for all six models tested (arXiv). KGMON's design goes after those two weak points: schema scouting targets joins, and pre-extracting prose and video targets cross-format evidence.

Limitations, safety and contested findings

  • NVIDIA says several choices fit the competition and may not suit production. Examples are extracting tables from documents and pre-processing every video.
  • Repeated attempts raise token use, latency and cost.
  • Automated self-improvement loops overfit the benchmark. The team saw training examples hardcoded into prompts, instructions that contradicted each other, and brittle post-processing rules.
Read the full section
  • NVIDIA says several choices fit the competition and may not suit production. Examples are extracting tables from documents and pre-processing every video.
  • Repeated attempts raise token use, latency and cost.
  • Automated self-improvement loops overfit the benchmark. The team saw training examples hardcoded into prompts, instructions that contradicted each other, and brittle post-processing rules. Its fix was held-out tasks, prompt audits and human approval before changes were kept (NVIDIA).
  • Without released code or ablations, KGMON's result cannot be reproduced.

Business and practitioner implications

  • Fix the plumbing first.
  • Treat traces as a product feature. Logging every tool call made failures diagnosable, including by a separate "inspector" agent that sorted failures into categories.
  • Use repeated attempts only when answers disagree or validation fails.
Read the full section
  • Fix the plumbing first. Teams on small open models may gain more from cleaner data access and fewer tools than from switching to a bigger model.
  • Treat traces as a product feature. Logging every tool call made failures diagnosable, including by a separate "inspector" agent that sorted failures into categories.
  • Spend extra compute selectively. Use repeated attempts only when answers disagree or validation fails.
  • Gate automated improvements. Agent-driven tuning needs held-out tests and human sign-off before changes are promoted.
  • Study the winners. The presentation archive and first-place code are the most concrete public references.
FOLLOW THE EVIDENCE

The source trail.

Sources (7)
A LITTLE LESS NOISE. A LOT MORE CONTEXT.

Stay curious.
Follow the evidence.

Independent perspectives, the original sources, and room for the questions that don't have easy answers.

How we build the brief
Connect with us

Find us where you already read.