[bot] Merge master/60203cf6 into rel/dev - #1848
Conversation
The Report copilot answers with a `report` multipart part carrying the drafted report as code. The SSE client did not list the type, so every report turn logged an unknown-part warning. It is now a known type and, like `dashboard`, stays in `unhandled_parts` verbatim for an evaluator to read back by type. jira: LX-3174 risk: nonprod Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…test The fixture invented a page shape. It now follows what gen-ai writes (composed_report.aac.json): format "widescreen" and a "column" layout. jira: LX-3174 risk: nonprod Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Scores whether asking the chat for a report returns one. The Report copilot keeps its draft in conversation state and saves nothing, so the evaluator reads the reply only: a successful draft_report call, a `report` part carrying the report document, its ref matching the draft's, a page count that agrees with the pages (a cover plus at least one content page), and a draft that is neither saved nor editing a saved report. A fixture may also state the period and the charts the report must show; those checks run only when it does. When the copilot asks back instead of drafting, a fixed reply built from the fixture answers it, as the dashboard skill does. Whether it asked first is recorded when the fixture expects a question, never gated: how much the copilot should ask is still an open product decision. Registered as agentic_report_skill; it runs serially until the dataset has runs behind it. jira: LX-3175 risk: nonprod Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A report fixture can now state a `narrative`: what the summaries must
cover. Two checks follow from it. report_summaries_present is
deterministic: the report's content pages have at least one `summary` slot, the slot
gen-ai writes its page summaries into, and every such slot carries
written text, with template placeholders such as {periodStart} not
counting as text. A content page laid out without a summary slot is
not a failure, and a static text slot is not a summary.
report_narrative_judged hands the report, rendered as plain text
(title, period, and per content page its heading, charts and summary),
to the binary LLM judge with the narrative as the expected output.
The judge is built only for a fixture that states a narrative, so the
other report items need neither the llm-judge extra nor
OPENAI_API_KEY. A run is ungraded only when the judge returned nothing
readable and the narrative was its one open check; a run that already
failed another check stays a failure. An ungraded run is left out of
pass@K, keeps pass^K from holding and writes no Langfuse scores, and
an item with no graded run raises JudgeResponseError, as the
general-question evaluator does.
jira: LX-3176
risk: nonprod
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
feat(gooddata-eval): evaluate the Report copilot
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## rel/dev #1848 +/- ##
===========================================
+ Coverage 83.84% 84.10% +0.25%
===========================================
Files 332 333 +1
Lines 22709 23099 +390
===========================================
+ Hits 19041 19428 +387
- Misses 3668 3671 +3 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
🚀 Automated PR to perform merge from master into rel/dev with changes up to 60203cf (created by https://github.com/gooddata/gooddata-python-sdk/actions/runs/37437684734).