Skip to content

Make gpt-5.6-luna usable: per-model temperature, json_schema graders - #186

Merged
adamjohnwright merged 4 commits into
mainfrom
feat/gpt-5.6-luna
Sep 9, 2026
Merged

adamjohnwright merged 4 commits into
mainfrom
feat/gpt-5.6-luna

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

Two things in this repo stopped gpt-5.6-luna working at all. Both fixed; the default model is unchanged — this makes the switch possible and tested, it does not make it.

1. Temperature

get_llm hardcoded temperature=0.0. The gpt-5.5/5.6 families and gpt-6-astra reject it:

Unsupported value: temperature does not support 0.0 with this model. Only the default (1) value is supported.

as a 400 on the first request, not at construction — so nothing notices until a user asks a question. resolve_temperature() sends 1.0 for those families, 0.0 for everything else, LLM_TEMPERATURE overrides.

The trap worth recording

Sending no temperature is not possible on langchain-openai 0.2.14: omitting the argument makes ChatOpenAI send its own pydantic default of 0.7, which these models refuse just as firmly as 0.0.

The first version of this branch did exactly that. model_dump(exclude_unset=True) agreed the field was unset — and the request still carried 0.7. Only running it against the live API showed it. That is constitution Article I in one line, so the test file says so.

What temperature 1 costs

Determinism. Measured, 10 runs each on four inputs:

input gpt-4o-mini @ 0.0 gpt-5.6-luna @ 1.0
science question → intent reactome ×10 reactome ×10
how-to question → intent userguide ×10 userguide ×10
benign text → safety true ×10 true ×10
prompt injection → safety false ×10 false ×10

Stable on what was tested. Not a guarantee.

2. Structured output

The three graders used the default function_calling method, which the gpt-5.6 family refuses on /v1/chat/completions:

Function tools with reasoning_effort are not supported for gpt-5.6-luna ... use /v1/responses or set reasoning_effort to none.

and langchain-openai 0.2.14 has no Responses API support. method="json_schema" uses response_format instead — verified for all three graders on both models, so this is not a luna-only path that would rot untested.

(reasoning_effort="none" also works, but it turns the reasoning off, which is the point of the model. With json_schema, every effort level works.)

End-to-end

Through AgentGraph.ainvoke on the React-to-Me profile — the same entry point bin/chat-chainlit.py uses — with the Release95 bundle and three real questions. Both models answer; 6 LLM calls each.

s/question input tok output tok
gpt-4o-mini 22.5 ~2685 see note
gpt-5.6-luna 41.2 ~3238 see note

Note: the OpenAI callback undercounts streamed completions inconsistently between the two, so treat output tokens as indicative only. Latency and call count are solid.

Pricing I cannot verify — no API reports it. That is the one input to the "just as cheap" question I could not check.

Qualitatively, on "Which complexes contain EGFR?" luna named four specific complexes with Reactome links; gpt-4o-mini gave a general description naming none. n=3, no rubric — a signal, not a measurement.

Two things in this repository stopped gpt-5.6-luna working at all. Both
are fixed here; the default model is unchanged.

1. Temperature. get_llm hardcoded temperature=0.0, which this repo wants
everywhere -- the graders, the intent classifier and the query expander
should give the same answer twice. The gpt-5.5/5.6 families and
gpt-6-astra reject it:

    Unsupported value: 'temperature' does not support 0.0 with this
    model. Only the default (1) value is supported.

as a 400 on the first request, not at construction, so nothing notices
until a user asks a question. resolve_temperature() sends 1.0 for those
families and 0.0 for everything else, with LLM_TEMPERATURE to override.

Sending *no* temperature is not an option, and this was the trap: on
langchain-openai 0.2.14, omitting the argument makes ChatOpenAI send its
own pydantic default of 0.7, which these models refuse just as firmly.
The first version of this commit did exactly that, and
model_dump(exclude_unset=True) confirmed the field was unset while the
request still carried 0.7. Only running it against the API showed it.
Hence a value, not an omission.

Temperature 1 costs determinism. Measured over 10 runs each on four
inputs -- science question, how-to question, benign text, prompt
injection -- both the intent classifier and the safety checker returned
the same verdict every time on gpt-5.6-luna as on gpt-4o-mini at 0.0.
Stable on what was tested; not a guarantee.

2. Structured output. The three graders used the default
function_calling method, which the gpt-5.6 family refuses on
/v1/chat/completions ("Function tools with reasoning_effort are not
supported ... use /v1/responses or set reasoning_effort to 'none'"), and
langchain-openai 0.2.14 has no Responses API support. method="json_schema"
uses response_format instead. Verified for all three graders against
both gpt-4o-mini and gpt-5.6-luna, so this is not a luna-only path that
would rot untested.

Verified end to end through AgentGraph.ainvoke on the React-to-Me
profile, the same entry point bin/chat-chainlit.py uses, with the
Release95 bundle and three real questions. Both models answer. Averaged
per question: gpt-4o-mini 22.5s, gpt-5.6-luna 41.2s, 6 LLM calls each.
@adamjohnwright
adamjohnwright merged commit 060d142 into main Sep 9, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the feat/gpt-5.6-luna branch September 9, 2026 14:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant