Skip to content

Add admin-configurable summary system prompt (SummaryProvider currently hardcodes it) #405

Description

@khnjrdm

Describe the feature you'd like to request

apps/integration_openai/lib/TaskProcessing/SummaryProvider.php currently hardcodes the system prompt used for text-summarisation tasks:

You are a helpful assistant that summarizes text in the same language as the text.
You should only return the summary without any additional information.

There is no admin UI or config:app:set key to override this prompt. The admin-visible Chat User Instructions for Chat Completions field applies only to the interactive chat-completion path, not to summarisation (which goes through TaskProcessing → SummaryProvider::process() on a separate code path). The max_tokens = 1000 default on this same code path is also very restrictive.

Combined, this produces very short generic summaries. Talk call-recording summaries in particular miss most of the value of the transcript — no structure, no attendees, no action items, no decisions, no next steps.

Every deployment that wants better summaries — richer meeting summaries, per-team wording, per-language variants, non-English house style, industry-specific vocabulary — has to patch PHP directly. That patch is clobbered by every app update and has to be manually re-applied. This is fragile across a fleet and is a real papercut for any admin running integration_openai against a LocalAI-compatible endpoint (LiteLLM, Ollama, self-hosted OpenAI-shape LLMs).

Environment where this was observed:

  • Nextcloud 33.0.6 (Hub 26 Winter)
  • integration_openai 4.5.1
  • Talk 23.0.8
  • Backend: LiteLLM proxy fronting local LLM (Qwen 2.5) — url config points at the LiteLLM endpoint

Describe the solution you'd like

Add a Summary system prompt text field to Admin → Connected accounts → OpenAI and LocalAI integration (or wherever fits the app's existing settings UI). The field should:

  • Persist to integration_openai app config under a stable key, e.g. summary_system_prompt
  • Default to the current hardcoded string, so upgrade is a no-op for existing deployments
  • Be read by SummaryProvider::process() in place of the hardcoded string; fall back to the current default when the config key is empty
  • Support multi-line input (textarea, not single-line)
  • Include short help text explaining the field is used verbatim as the system message

Sketch of the code change in SummaryProvider.php:

// before
$systemPrompt = 'You are a helpful assistant that summarizes text in the same language as the text. '
              . 'You should only return the summary without any additional information.';

// after
$default = 'You are a helpful assistant that summarizes text in the same language as the text. '
         . 'You should only return the summary without any additional information.';
$systemPrompt = $this->appConfig->getAppValueString('summary_system_prompt', $default) ?: $default;

Code Change 2 (related, same file): consider also surfacing max_tokens for summarisation as a UI field. It's already tunable via occ config:app:set integration_openai max_tokens --value=8000, but the pairing (prompt + token budget) is what actually determines summary quality — surfacing both in the same UI section makes them discoverable together.

Broader shape (nice-to-have, out of scope for this ticket): the same "hardcoded prompt with no override" pattern likely applies to other core:text2text:* task types (headline, chat-summary, etc.). A broader "per-task-type prompt templates" epic would be the ideal end state, but even just the Summary field would unblock the most common request today.

Describe alternatives you've considered

  1. Patching SummaryProvider.php directly on each server. Works, but the patch is clobbered on every app upgrade. Not maintainable across a fleet, and every admin who wants a better summary has to re-derive the same patch.

  2. Using the existing "Chat User Instructions for Chat Completions" field. That field only applies to the interactive chat path, not to TaskProcessing summarisation. Confirmed by inspecting SummaryProvider::process() — the chat instructions config is not consulted there.

  3. Using a middleware proxy (LiteLLM, LocalAI itself, etc.) to rewrite the system prompt. LiteLLM can inject or replace system prompts on a per-model basis. This works, but pushes app-level configuration into infrastructure, is invisible to Nextcloud admins, and means every deployment needs a proxy in front of the LLM endpoint just for this one setting.

  4. Waiting for a broader "per-task-type prompt templates" feature. Ideal end state, but shipping just the Summary override in the meantime would unblock most of the actual pain today.

Activity

  1. lukasdotcom commented on Jul 17, 2026

    @lukasdotcom
    Member

    Hi @khnjrdm,

    For the max tokens you can it in the UI it is hidden under advanced input options.
    Image

    Allowing admins to modify the prompt for the providers is a bit out of scope for this app and would result in hard to troubleshoot configurations. Also the next release is bringing out some more customization options for summarize (#387) and would be incompatible with allowing admins to change the prompt.

  2. khnjrdm commented on Jul 17, 2026

    @khnjrdm
    Author

    Thanks @lukasdotcom, and thanks for the screenshot — that clarifies which UI you meant.

    On max_tokens: the "advanced input options" you're showing is the Assistant Summarize task UI, where a admin is actively invoking summarize and can set max words/tokens per call. That's distinct from the app-level default I actually had to change to make Talk recording summaries usable, which is Admin → Connected accounts → Text generation → Max output tokens per request = 8000:

    Image

    Confirming which of these the automated Talk flow actually reads would be useful either way, but that's not the core of this ticket.

    The core of the ticket is the automated Talk recording → summary flow, and I don't think #387 or the advanced input options reach it. When someone finishes recording a Talk call, spreed queues a core:text2text:summary TaskProcessing task from server-side code — there is no human sitting in front of the Assistant Summarize UI at that moment to drag a max words/tokens slider or a #387 complexity/length knob. The task fires with defaults; SummaryProvider::process() reads the app-level config, uses the hardcoded prompt, and writes Recording <ts> - summary.md into the recording folder.

    What the default prompt produces on that automated path (verbatim from SummaryProvider.php):

    You are a helpful assistant that summarizes text in the same language as the text.
    You should only return the summary without any additional information.
    

    On a 60-minute meeting transcript (LiteLLM + Qwen 2.5 7B, also verified against GPT-4o-mini) this produces 2–4 sentences of generic narration, "The team discussed several topics including X and Y", with no attendees, no decisions, no action items, no next steps. It's the shape of a book blurb, not meeting minutes. For its actual use (internal recap someone reads the next morning), that's not useful.

    What we're running in production after patching SummaryProvider.php:

    You are analyzing a transcript, typically of a business meeting, workshop, or client call.
    Produce a structured summary in Markdown in the SAME LANGUAGE as the source text.
    Detect the language automatically; do not mention which language you detected.
    Include these sections when the transcript supports them:
    **Meeting overview** (2-3 sentences of context);
    **Attendees** (if identifiable);
    **Key topics discussed** (bullet points, 1-2 sentences each);
    **Decisions made** (concrete outcomes);
    **Action items** (who committed to what, by when if stated);
    **Open questions** (items requiring follow-up);
    **Next steps** (planned follow-ups or milestones).
    Omit sections that have no content. Do not fabricate action items, decisions, or attendees not present in the transcript.
    For non-meeting content or very short input, produce a proportionally shorter summary without imposing this structure.
    

    Same model, same transcript, this produces structured Markdown minutes with real attendees, decisions, and action items. The delta isn't stylistic, it's structural. Complexity + length are shape modifiers (how detailed, how long); they can't request a specific document structure. LLMs don't put attendees / decisions / action items into distinct sections unless the prompt asks for it, and a "detailed, long" summary of a meeting is still narrative rather than minutes.

    Three smaller-scope options that sidestep the "admin can rewrite any prompt is a foot-gun" concern, any of these solves the automated Talk flow:

    1. occ-only config, no UI: integration_openai.summary_system_prompt via config:app:set only. Default = today's string; admins who override it have explicitly taken the risk. Same pattern as spreed.call_recording_transcription.
    2. Move meeting-summary out of the generic path: let spreed (Talk) pass a prompt hint in the TaskProcessing input when the summary source is a call recording. The caller knows it's a meeting so it asks for meeting-shaped output. integration_openai stays generic.
    3. New task subtype: core:text2text:recording-summary with a meeting-oriented default in core. Bigger change but solves "generic summary is wrong for meetings" across the stack.

    Happy to run comparative output on our stack (NC 33.0.6 + Talk 23.0.8 + LiteLLM + Qwen 2.5) against real recordings if that helps decide. Otherwise fine to close as won't-fix and keep patching locally, just wanted to make sure the ticket wasn't read as "let admins rewrite any prompt they want" when the pain is specifically automated Talk recording summaries.

  3. Abhijeet-035 commented on Sep 6, 2026

    @Abhijeet-035
    Contributor

    Hi @khnjrdm ! I’d like to work on this issue.

    I checked the current implementation. SummaryProvider still has the summary system prompt hardcoded, while max_tokens is already handled by the current Task Processing implementation (SummaryProvider.php).

    I’m planning to:

    • add an admin-configurable summary_system_prompt setting, using the current hardcoded prompt as the default.
    • expose it as a multiline textarea in the admin settings.
    • use the configured prompt in SummaryProvider, with a fallback to the default when empty.
    • preserve the existing format and complexity instructions.

    Please let me know if this approach looks good before I start working on the PR.

  4. khnjrdm commented on Sep 10, 2026

    @khnjrdm
    Author

    @Abhijeet-035 Yes, from my perspective that would make things easier to deal with and my own patches wouldn't be needed even if they've so far survived updates of Nextcloud.

  5. Abhijeet-035 commented on Sep 10, 2026

    @Abhijeet-035
    Contributor

    Thanks for confirming, @khnjrdm I've implemented the proposed approach and opened PR #433. It adds an admin-configurable summary system prompt, falls back to the existing default when empty, and keeps the existing format and complexity instructions.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions