feat(gooddata-eval): add the agentic report-skill evaluator - #1844
Merged
Merged
Conversation
Scores whether asking the chat for a report returns one. The Report copilot keeps its draft in conversation state and saves nothing, so the evaluator reads the reply only: a successful draft_report call, a `report` part carrying the report document, its ref matching the draft's, a page count that agrees with the pages (a cover plus at least one content page), and a draft that is neither saved nor editing a saved report. A fixture may also state the period and the charts the report must show; those checks run only when it does. When the copilot asks back instead of drafting, a fixed reply built from the fixture answers it, as the dashboard skill does. Whether it asked first is recorded when the fixture expects a question, never gated: how much the copilot should ask is still an open product decision. Registered as agentic_report_skill; it runs serially until the dataset has runs behind it. jira: LX-3175 risk: nonprod Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
Contributor
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
agentic_report_skill, a gd-eval evaluator that checks whether asking the chat for a report actually returns one. It scores the reply only, since the Report copilot saves nothing; the LLM-judge score for the narrative is a separate PR.LX-3175, part of LX-3166. Stacked on #1841 (base
rr/LX-3174-report-part), which teaches the SSE client thereportpart.What this implements
A fixture states what the report must show; every key is optional:
{"id": "report-vague-prompt", "test_kind": "agentic_report_skill", "question": "Make me a report.", "expected_output": {"period": {"start": "2026-01-01", "end": "2026-06-30"}, "expects_clarification": true}}Run locally against
lynx-agentson staging (demoworkspace, gpt-5.2,--runs 1):Scored checks (
report_prefix socombo_report.pycan attribute them):report_drafteddraft_reportcall succeeded (the last successful one counts)report_part_presentreportpart carries a non-null document withtype == "report"report_ref_matchesreport_refis therefthe draft returned, and is not emptyreport_pages_consistentlen(pages), the part'spage_countand the tool'spage_countagree, and there are at least 2 pages (cover + content)report_not_savedsaved_report_idandbase_report_idare both nullreport_skill_activatedset_skillsactivatedreport_builder(a run with no routing call passes, as in the dashboard skill)report_period_correctperiodreport_charts_matchedvisualizations; ids are found anywhere in the nestedcolumn/rowlayoutDecisions
How much the copilot should ask before drafting is still an open product decision, so a run that drafts straight away must not fail.
report_asked_firstis published only when the fixture setsexpects_clarification, and it reads turn 1 throughclassify_reply, so a refusal followed by a draft does not count as asking.quality_score(88% above) the same way the dashboard skill's diagnostics do. I'd accept moving it out of the detail if that reads as a penalty.Same as the dashboard skill: a failure stays the copilot's. Up to 4 turns; a silent turn ends the run.
build_simulated_replyincore/agentic/report_skill.pydraft_reportresults reach the stream as plainmodel_dump_json()(nodatawrapper); the part'spage_countandbase_report_idcome from the same stored draft as the tool result, so agreement is the correct expectation; an unresolved part arrives as an emptyReportPart, which failsreport_part_present.agentic_report_skillruns serially.It stays off
PARALLEL_SAFE_TEST_KINDSlike the dashboard skill, until the dataset has runs behind it.cli/agentic_runner.py,core/agentic/__init__.pyTest plan
tests/test_agentic_report_skill.py: 38 tests. Scoring is pure and tested branch by branch; the conversation loop runs against a scripted fakeChatClient. Sabotaging the ref check, the minimum-pages check andasked_firsteach turned a test red.What comes next
risk: nonprod
🤖 Generated with Claude Code