Skip to content

Switch label_misinformation to Mistral by default - #125

Open
gmguarino wants to merge 8 commits into
mainfrom
feat/update-model-to-mistral
Open

gmguarino wants to merge 8 commits into
mainfrom
feat/update-model-to-mistral

Conversation

@gmguarino

@gmguarino gmguarino commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Adds MistralSinglePromptPipeline to jobs/label_misinformation alongside the existing OpenAI SinglePromptPipeline, selectable via a new LLM_PROVIDER env var (default: mistral). Uses mistral-small-2603 with a new French structured prompt (reasoning + misinformation bool, prompt version "0.1.0"), hardened against a repetition-loop failure mode found during evaluation (penalties, stop sequence, retry-on-parse-failure).
  • Whether MIN_MISINFORMATION_SCORE thresholding applies is now a property of the active prompt (DisinformationPrompt.binary), not the provider — binary prompts (the new Mistral one) filter on score > 0 directly; continuous prompts (the existing OpenAI one) keep using the threshold.
  • Adds a standalone OpenAI eval script (climateguard/finetuning/api/openai/completions/test.py) mirroring the existing Mistral one, used to compare the two providers on the DataForGood/climateguard train/test splits before choosing Mistral.
  • Fixes an "Event loop is closed" bug in the Mistral eval script (client was reused across separate asyncio.run() calls).
  • Documents all of the above in jobs/README.md.

Compatibility / rollout checklist

For this repo/branch to work as-is (Mistral, the new default):

  • Create jobs/label_misinformation/secrets/mistral_key locally (and the equivalent secret in whatever deploys this job — Vaultwarden/CI/Scaleway) with a real Mistral API key. Not included in this PR since it's a secret.
  • Confirm MISTRAL_API_KEY is available wherever docker-compose.yml's app/test/testconsole services run (the mistral_key docker secret is already wired up in this PR).
  • Run docker compose up testconsole -d && docker compose exec testconsole pytest tests/test_pipeline.py tests/test_prompts.py -vv to sanity-check the new pure-function tests pass.
  • Smoke-test docker compose up app against a small known day/channel and confirm in the logs it used the Mistral provider, and that flagged rows have sensible French model_reason text.

For any job/service that should keep using OpenAI models (nothing removed — OpenAI support is fully intact as an opt-in path):

  • Set LLM_PROVIDER=openai explicitly in that service's environment (main.py defaults to mistral otherwise).
  • Ensure OPENAI_API_KEY / the openai_key docker secret is still populated (unchanged from before this PR).
  • No prompt config needed — the OpenAI path always uses prompt version "0.0.1" (unchanged, English, plain 0-10 score) regardless of PIPELINE_PRODUCTION_PROMPT.
  • MIN_MISINFORMATION_SCORE continues to apply exactly as before for this path (untouched behavior).
  • MODEL_NAME/MODEL_NAME_BRAZIL should be set to an OpenAI model name (e.g. gpt-4o-mini) for that service, since country.py's defaults were changed to mistral-small-2603 to match the new default provider.

Out of scope (left untouched, per explicit request): the shadow docker-compose service and requirements_shadow.txt (BERT shadow pipeline) — still OpenAI-only, no LLM_PROVIDER wiring added there.

Test plan

  • parse_structured_response unit tests (valid JSON, malformed-JSON-with-regex-fallback, unparseable-defaults-to-zero) run locally and pass.
  • Existing parse_response/parse_response_reason tests re-verified locally, still pass.
  • test_prompts.py's existing assertions re-verified against the new default (PIPELINE_PRODUCTION_PROMPT now resolves to "0.1.0", prod=True).
  • All touched files pass py_compile.
  • Full docker compose up test / app smoke test (needs live secrets/DB — not run in this environment).

Creating a fresh client outside asyncio.run() while calling asyncio.run()
separately per split closed the event loop the client's connections were
bound to, causing "Event loop is closed" on the second split. Also bumps
MAX_CONCURRENCY to 20.
Mirrors the mistral eval script (same climateguard/climate misinformation
dataset, train/test splits, concurrency, retry-on-parse-failure) but calls
OpenAI's chat completions API with production's plain-number prompt, and
binarizes the 0-10 score against MIN_MISINFORMATION_SCORE before scoring, to
compare against the Mistral structured-output approach on equal footing.
Adds MistralSinglePromptPipeline (structured JSON output: reasoning +
misinformation bool, hardened against the repetition-loop failure mode
found during eval) alongside the existing OpenAI SinglePromptPipeline,
selectable per-job via LLM_PROVIDER (default: mistral). Whether
MIN_MISINFORMATION_SCORE thresholding applies is now a property of the
active prompt (DisinformationPrompt.binary) rather than the provider,
since that's what actually determines whether the model's score is
continuous (0-10) or a binary flag (0 or 10).

Also updates country.py's default model to mistral-small-2603, adds the
mistralai dependency, wires MISTRAL_API_KEY/LLM_PROVIDER through
docker-compose for the app/test/testconsole services (shadow untouched),
and documents the new configuration in jobs/README.md.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant