Skip to content

rubric_based_final_response_quality_v1 errors when the agent replies from a before_agent_callback聽#7379

Description

@MihailRussu

馃敶 Required Information

Describe the Bug:
Since 2.10, an eval case whose agent answers from a before_agent_callback (so no model call) can no longer be judged with rubric_based_final_response_quality_v1. The metric fails with `my_agent` not found in the agentic system. and the case comes back NOT_EVALUATED, although the agent's reply is the same as on 2.9.

Steps to Reproduce:

  1. Define an LlmAgent whose before_agent_callback returns a types.Content, e.g. "Authentication required."
  2. Add an eval case for it with any rubric under rubric_based_final_response_quality_v1.
  3. Run it with adk eval (or LocalEvalService).

Expected Behavior:
The rubric is judged against the reply, as on 2.9.0.

Observed Behavior:
local_eval_service logs Metric evaluation failed for metric rubric_based_final_response_quality_v1 ... with following error my_agent not found in the agentic system. and the case is NOT_EVALUATED on every run.

Environment Details:

  • ADK Library Version: 2.11.0 (also 2.10.0; 2.9.0 works)
  • Desktop OS: Linux
  • Python Version: 3.13

Model Information:

  • Are you using LiteLLM: No
  • Which model is being used: the agent makes no model call; the judge is a Gemini model

馃煛 Optional Information

Regression:
Yes, works on 2.9.0.

Additional Context:
As far as I can tell this comes from 2d428ea (informational efficiency metrics), which stopped dropping the final event from invocation_events (it is now kept with content=None to carry usage metadata). rubric_based_final_response_quality_v1 takes invocation_events[0].author as the agent name and calls AppDetails.get_developer_instructions(), which raises because app details are only recorded from intercepted model requests, and a callback short-circuit never makes one. On 2.9 the list was empty, so the judge fell back to next(iter(app_details.agent_details)), found nothing and went on without developer instructions.

Falling back to empty developer instructions when the author is not in agent_details (or only taking the author when it is) would restore the 2.9 behaviour. hallucinations_v1 and the multi-turn trajectory evaluator look fine, since they iterate the recorded agents or only use the name as a label.

Minimal Reproduction Code:

from google.adk.agents import Agent
from google.genai import types


def require_token(callback_context):
    if not callback_context.state.get('access_token'):
        return types.Content(role='model', parts=[types.Part(text='Authentication required.')])
    return None


root_agent = Agent(
    name='my_agent',
    model='gemini-2.5-flash',
    instruction='Answer questions about the user account.',
    before_agent_callback=require_token,
)

with an eval case whose session state has no access_token and this criterion:

{"criteria": {"rubric_based_final_response_quality_v1": {"threshold": 0.6, "rubrics": [
  {"rubric_id": "refuses", "rubric_content": {"text_property": "The agent asks the user to authenticate."}}]}}}

How often has this issue occurred?:

  • Always (100%)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

eval[Component] This issue is related to evaluation

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions