Repository navigation
Conversation
Creating a fresh client outside asyncio.run() while calling asyncio.run() separately per split closed the event loop the client's connections were bound to, causing "Event loop is closed" on the second split. Also bumps MAX_CONCURRENCY to 20.
Mirrors the mistral eval script (same climateguard/climate misinformation dataset, train/test splits, concurrency, retry-on-parse-failure) but calls OpenAI's chat completions API with production's plain-number prompt, and binarizes the 0-10 score against MIN_MISINFORMATION_SCORE before scoring, to compare against the Mistral structured-output approach on equal footing.
Adds MistralSinglePromptPipeline (structured JSON output: reasoning + misinformation bool, hardened against the repetition-loop failure mode found during eval) alongside the existing OpenAI SinglePromptPipeline, selectable per-job via LLM_PROVIDER (default: mistral). Whether MIN_MISINFORMATION_SCORE thresholding applies is now a property of the active prompt (DisinformationPrompt.binary) rather than the provider, since that's what actually determines whether the model's score is continuous (0-10) or a binary flag (0 or 10). Also updates country.py's default model to mistral-small-2603, adds the mistralai dependency, wires MISTRAL_API_KEY/LLM_PROVIDER through docker-compose for the app/test/testconsole services (shadow untouched), and documents the new configuration in jobs/README.md.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
MistralSinglePromptPipelinetojobs/label_misinformationalongside the existing OpenAISinglePromptPipeline, selectable via a newLLM_PROVIDERenv var (default:mistral). Usesmistral-small-2603with a new French structured prompt (reasoning+misinformationbool, prompt version"0.1.0"), hardened against a repetition-loop failure mode found during evaluation (penalties, stop sequence, retry-on-parse-failure).MIN_MISINFORMATION_SCOREthresholding applies is now a property of the active prompt (DisinformationPrompt.binary), not the provider — binary prompts (the new Mistral one) filter onscore > 0directly; continuous prompts (the existing OpenAI one) keep using the threshold.climateguard/finetuning/api/openai/completions/test.py) mirroring the existing Mistral one, used to compare the two providers on theDataForGood/climateguardtrain/test splits before choosing Mistral.asyncio.run()calls).jobs/README.md.Compatibility / rollout checklist
For this repo/branch to work as-is (Mistral, the new default):
jobs/label_misinformation/secrets/mistral_keylocally (and the equivalent secret in whatever deploys this job — Vaultwarden/CI/Scaleway) with a real Mistral API key. Not included in this PR since it's a secret.MISTRAL_API_KEYis available whereverdocker-compose.yml'sapp/test/testconsoleservices run (themistral_keydocker secret is already wired up in this PR).docker compose up testconsole -d && docker compose exec testconsole pytest tests/test_pipeline.py tests/test_prompts.py -vvto sanity-check the new pure-function tests pass.docker compose up appagainst a small known day/channel and confirm in the logs it used the Mistral provider, and that flagged rows have sensible Frenchmodel_reasontext.For any job/service that should keep using OpenAI models (nothing removed — OpenAI support is fully intact as an opt-in path):
LLM_PROVIDER=openaiexplicitly in that service's environment (main.py defaults tomistralotherwise).OPENAI_API_KEY/ theopenai_keydocker secret is still populated (unchanged from before this PR)."0.0.1"(unchanged, English, plain 0-10 score) regardless ofPIPELINE_PRODUCTION_PROMPT.MIN_MISINFORMATION_SCOREcontinues to apply exactly as before for this path (untouched behavior).MODEL_NAME/MODEL_NAME_BRAZILshould be set to an OpenAI model name (e.g.gpt-4o-mini) for that service, sincecountry.py's defaults were changed tomistral-small-2603to match the new default provider.Out of scope (left untouched, per explicit request): the
shadowdocker-compose service andrequirements_shadow.txt(BERT shadow pipeline) — still OpenAI-only, noLLM_PROVIDERwiring added there.Test plan
parse_structured_responseunit tests (valid JSON, malformed-JSON-with-regex-fallback, unparseable-defaults-to-zero) run locally and pass.parse_response/parse_response_reasontests re-verified locally, still pass.test_prompts.py's existing assertions re-verified against the new default (PIPELINE_PRODUCTION_PROMPTnow resolves to"0.1.0",prod=True).py_compile.docker compose up test/appsmoke test (needs live secrets/DB — not run in this environment).