Skip to content

feat(gooddata-eval): add the agentic report-skill evaluator - #1844

Merged
romrak merged 1 commit into
rr/LX-3174-report-partfrom
rr/LX-3175-report-skill
Oct 5, 2026
Merged

romrak merged 1 commit into
rr/LX-3174-report-partfrom
rr/LX-3175-report-skill

Conversation

@romrak

@romrak romrak commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Adds agentic_report_skill, a gd-eval evaluator that checks whether asking the chat for a report actually returns one. It scores the reply only, since the Report copilot saves nothing; the LLM-judge score for the narrative is a separate PR.

LX-3175, part of LX-3166. Stacked on #1841 (base rr/LX-3174-report-part), which teaches the SSE client the report part.

What this implements

A fixture states what the report must show; every key is optional:

{"id": "report-vague-prompt", "test_kind": "agentic_report_skill", "question": "Make me a report.",
 "expected_output": {"period": {"start": "2026-01-01", "end": "2026-06-30"}, "expects_clarification": true}}

Run locally against lynx-agents on staging (demo workspace, gpt-5.2, --runs 1):

report-clear-prompt  PASS  35.82s  quality=100%
  report_drafted, report_part_present, report_ref_matches, report_pages_consistent,
  report_not_saved, report_skill_activated, report_period_correct: all True; summaries_from_data=2
report-vague-prompt  PASS  49.18s  quality=88%
  the same checks all True; report_asked_first=False (it drafted with the default
  half-year instead of asking), summaries_from_data=3

Scored checks (report_ prefix so combo_report.py can attribute them):

Score Passes when
report_drafted a draft_report call succeeded (the last successful one counts)
report_part_present a report part carries a non-null document with type == "report"
report_ref_matches the part's report_ref is the ref the draft returned, and is not empty
report_pages_consistent len(pages), the part's page_count and the tool's page_count agree, and there are at least 2 pages (cover + content)
report_not_saved saved_report_id and base_report_id are both null
report_skill_activated set_skills activated report_builder (a run with no routing call passes, as in the dashboard skill)
report_period_correct only if the fixture states period
report_charts_matched only if the fixture lists visualizations; ids are found anywhere in the nested column/row layout

Decisions

  1. Asking first is recorded, never gated.
    How much the copilot should ask before drafting is still an open product decision, so a run that drafts straight away must not fail. report_asked_first is published only when the fixture sets expects_clarification, and it reads turn 1 through classify_reply, so a refusal followed by a draft does not count as asking.
  • It is a boolean in the detail, so it lowers quality_score (88% above) the same way the dashboard skill's diagnostics do. I'd accept moving it out of the detail if that reads as a penalty.
  1. The simulated user's reply is fixed, built from the fixture, not written by an LLM.
    Same as the dashboard skill: a failure stays the copilot's. Up to 4 turns; a silent turn ends the run.
  • build_simulated_reply in core/agentic/report_skill.py
  1. The wire shapes are checked against gen-ai, not assumed.
    draft_report results reach the stream as plain model_dump_json() (no data wrapper); the part's page_count and base_report_id come from the same stored draft as the tool result, so agreement is the correct expectation; an unresolved part arrives as an empty ReportPart, which fails report_part_present.
  2. agentic_report_skill runs serially.
    It stays off PARALLEL_SAFE_TEST_KINDS like the dashboard skill, until the dataset has runs behind it.
  • cli/agentic_runner.py, core/agentic/__init__.py

Test plan

  • tests/test_agentic_report_skill.py: 38 tests. Scoring is pure and tested branch by branch; the conversation loop runs against a scripted fake ChatClient. Sabotaging the ref check, the minimum-pages check and asked_first each turned a test red.
  • Full package suite: 1361 passed. ruff, ruff format and ty clean. CodeRabbit: no findings.
  • Live run above.

What comes next

  • LX-3176: the LLM-judge narrative score, then a gooddata-eval release.
  • LX-3177: the Tavern wiring in gdc-nas.

risk: nonprod

🤖 Generated with Claude Code

Scores whether asking the chat for a report returns one. The Report
copilot keeps its draft in conversation state and saves nothing, so the
evaluator reads the reply only: a successful draft_report call, a
`report` part carrying the report document, its ref matching the
draft's, a page count that agrees with the pages (a cover plus at least
one content page), and a draft that is neither saved nor editing a
saved report. A fixture may also state the period and the charts the
report must show; those checks run only when it does.

When the copilot asks back instead of drafting, a fixed reply built
from the fixture answers it, as the dashboard skill does. Whether it
asked first is recorded when the fixture expects a question, never
gated: how much the copilot should ask is still an open product
decision. Registered as agentic_report_skill; it runs serially until
the dataset has runs behind it.

jira: LX-3175
risk: nonprod

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@romrak
romrak requested review from hkad98, lupko and pcerny as code owners October 5, 2026 11:29
@coderabbitai

coderabbitai Bot commented Oct 5, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Essentials
  • Run ID: b9acf001-a615-4ae6-ac7a-9f6a37ad0445

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@romrak
romrak merged commit 3594337 into rr/LX-3174-report-part Oct 5, 2026
1 check passed
@romrak
romrak deleted the rr/LX-3175-report-skill branch October 5, 2026 11:32
@romrak

romrak commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Folded into #1841: its commit was pushed onto #1841's branch, which is why GitHub shows this PR as merged. Nothing reached master; review it on #1841, which now carries LX-3174, LX-3175 and LX-3176.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant