Make gpt-5.6-luna usable: per-model temperature, json_schema graders - #186
Merged
Merged
Conversation
Two things in this repository stopped gpt-5.6-luna working at all. Both
are fixed here; the default model is unchanged.
1. Temperature. get_llm hardcoded temperature=0.0, which this repo wants
everywhere -- the graders, the intent classifier and the query expander
should give the same answer twice. The gpt-5.5/5.6 families and
gpt-6-astra reject it:
Unsupported value: 'temperature' does not support 0.0 with this
model. Only the default (1) value is supported.
as a 400 on the first request, not at construction, so nothing notices
until a user asks a question. resolve_temperature() sends 1.0 for those
families and 0.0 for everything else, with LLM_TEMPERATURE to override.
Sending *no* temperature is not an option, and this was the trap: on
langchain-openai 0.2.14, omitting the argument makes ChatOpenAI send its
own pydantic default of 0.7, which these models refuse just as firmly.
The first version of this commit did exactly that, and
model_dump(exclude_unset=True) confirmed the field was unset while the
request still carried 0.7. Only running it against the API showed it.
Hence a value, not an omission.
Temperature 1 costs determinism. Measured over 10 runs each on four
inputs -- science question, how-to question, benign text, prompt
injection -- both the intent classifier and the safety checker returned
the same verdict every time on gpt-5.6-luna as on gpt-4o-mini at 0.0.
Stable on what was tested; not a guarantee.
2. Structured output. The three graders used the default
function_calling method, which the gpt-5.6 family refuses on
/v1/chat/completions ("Function tools with reasoning_effort are not
supported ... use /v1/responses or set reasoning_effort to 'none'"), and
langchain-openai 0.2.14 has no Responses API support. method="json_schema"
uses response_format instead. Verified for all three graders against
both gpt-4o-mini and gpt-5.6-luna, so this is not a luna-only path that
would rot untested.
Verified end to end through AgentGraph.ainvoke on the React-to-Me
profile, the same entry point bin/chat-chainlit.py uses, with the
Release95 bundle and three real questions. Both models answer. Averaged
per question: gpt-4o-mini 22.5s, gpt-5.6-luna 41.2s, 6 LLM calls each.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two things in this repo stopped
gpt-5.6-lunaworking at all. Both fixed; the default model is unchanged — this makes the switch possible and tested, it does not make it.1. Temperature
get_llmhardcodedtemperature=0.0. The gpt-5.5/5.6 families andgpt-6-astrareject it:as a 400 on the first request, not at construction — so nothing notices until a user asks a question.
resolve_temperature()sends1.0for those families,0.0for everything else,LLM_TEMPERATUREoverrides.The trap worth recording
Sending no temperature is not possible on langchain-openai 0.2.14: omitting the argument makes
ChatOpenAIsend its own pydantic default of 0.7, which these models refuse just as firmly as 0.0.The first version of this branch did exactly that.
model_dump(exclude_unset=True)agreed the field was unset — and the request still carried 0.7. Only running it against the live API showed it. That is constitution Article I in one line, so the test file says so.What temperature 1 costs
Determinism. Measured, 10 runs each on four inputs:
reactome×10reactome×10userguide×10userguide×10true×10true×10false×10false×10Stable on what was tested. Not a guarantee.
2. Structured output
The three graders used the default
function_callingmethod, which the gpt-5.6 family refuses on/v1/chat/completions:and langchain-openai 0.2.14 has no Responses API support.
method="json_schema"usesresponse_formatinstead — verified for all three graders on both models, so this is not a luna-only path that would rot untested.(
reasoning_effort="none"also works, but it turns the reasoning off, which is the point of the model. Withjson_schema, every effort level works.)End-to-end
Through
AgentGraph.ainvokeon the React-to-Me profile — the same entry pointbin/chat-chainlit.pyuses — with the Release95 bundle and three real questions. Both models answer; 6 LLM calls each.Note: the OpenAI callback undercounts streamed completions inconsistently between the two, so treat output tokens as indicative only. Latency and call count are solid.
Pricing I cannot verify — no API reports it. That is the one input to the "just as cheap" question I could not check.
Qualitatively, on "Which complexes contain EGFR?" luna named four specific complexes with Reactome links; gpt-4o-mini gave a general description naming none. n=3, no rubric — a signal, not a measurement.