ACL 2026 Main

SituW situation working memory

Injecting Context via Situation Working Memory for Logical Reasoning with LLMs

Jieun Kim·Seoha Lim·YoungHae Choi·Sung-Bae Cho

Yonsei University

SituW overview: the premise is split into sentences; for each sentence a prompt extracts new situational elements, which update the SituW memory of time, space, protagonist, intention and causality; final reasoning combines each sentence with its memory to decide entailment.
The context is split into sentences and read in order. For each sentence, an LLM extracts newly implied situational elements and updates the situation working memory, which then guides the final reasoning.

motivation

Reasoning needs to know what is true now

Reading “Peter took the elevator to the fifth floor. He went to talk to his professor,” people infer that the professor’s office is probably on the fifth floor. Cognitive psychology calls this mental representation a situation model. It is organized along time, space, causality, intention and protagonist, and it is updated with every sentence.

LLM logical reasoning often converts text into formal logic instead. Translation loses contextual nuance, and symbolic solvers break on imprecise expressions. SituW gives the LLM an explicit, evolving situation memory to reason over.

Human reasoning tracks time, space, intention, causality and protagonists for each sentence of the Peter story; SituW reasoning passes the context through objective and subjective indexing into a situation working memory used by the LLM.

abstract

An explicit situation model for LLM reasoning

Recent advances in large language models (LLMs) have improved logical reasoning by incorporating formal logic or explicit structured representations. However, such methods often lose track of what is true now in multi-step reasoning, failing to maintain a coherent global state and its logical consequences. Motivated by Situation Model Theory in cognitive psychology, which views comprehension as constructing and updating a mental model of events along key dimensions (time, space, causality, intention, protagonist), we propose a cognitively inspired method of Situation Working Memory (SituW) for contextual reasoning in LLMs. SituW first builds a situation representation by decomposing text along these five dimensions, and guides LLM inference with the evolving state. Keeping an explicit, dynamically updated situation memory instead of a static logical form encourages globally consistent reasoning over the situation model rather than raw text. Evaluated in both supervised and prompt-based settings, SituW improves accuracy by 23.3%p and 15.93%p while reducing “uncertain” predictions, suggesting that explicit situation modeling supports more globally consistent LLM reasoning.

probe

Do LLMs build situation models?

We adapt a classic mental-model experiment to text. Each trial has four sentences that differ in one preposition, and the model picks the confusable pair: the two sentences most likely to describe the same event.

example item
  1. The girl was given a complete pedicure at the chiropodist’s.
  2. The girl was given a complete pedicure by the chiropodist.
  3. The girl had her handbag stolen at the chiropodist’s.
  4. The girl had her handbag stolen by the chiropodist.

Eleven graduate students reach 89.28% on 24 items. LLMs fall far below and seem to rely on surface overlap. Prompting GPT models to first extract time, location, protagonists, cause and intention raises accuracy by 20.8 points for GPT-3.5 and 8.3 for GPT-4o-mini.

Confusable pair accuracy: GPT-3.5 and GPT-4o improve by 20.8 and 8.3 points with situation indexing; open LLMs Llama3 8B, Llama3 70B, Llama3.3 70B and Mistral 7B stay well below the human line at 89.28%.

method

Build the memory sentence by sentence, then reason over it

The situation memory has two parts. Objective indexing records what the text states: time, space and protagonists. Subjective indexing infers what it implies: causality and intention.

context Zhu Hong: red squirrels make holes in the bark of sugar pines to absorb sap. Since the sap of sugar pine is mainly composed of water and a small amount of sugar, it is roughly certain that red squirrels are looking for water or sugar. Water is easily available in other ways where pine trees grow. Therefore, red pine trees are not trying to dig holes because they are looking for water, they may be looking for sugar. Lina: it must not be looking for sugar but something else, because the concentration of sugar in sugar pine sap is so low that red squirrels have to drink a lot of sap to get a little sugar.

01Objective indexing

Time
(None)
Space
sugar pines
Protagonist
Lina, Red squirrels
02 read memory →

03Subjective indexing

Causality
The low concentration of sugar in sugar pine sap
Intention
To absorb sap from sugar pines

Situation Working Memory

objectiveTime(None)
Spacesugar pines
ProtagonistZhu Hong, Red squirrels, Lina
subjectiveCausalityRed squirrels searching for water or sugar; the low concentration of sugar in sugar pine sap
IntentionSearching for water or sugar; to absorb sap from sugar pines
S₁ ⊕ m₁+S₂ ⊕ m₂+ … +question →LLM: plan · action · update · process →True / False / Uncertain

Example from Figure 7 of the paper. The updated memory adds the new protagonist and cause to the stored state.

Prompt-based construction

At inference time, one short prompt per element fills the memory. Each memory element is then verbalized and appended to the context, objective elements first. The LLM reasons over the enriched context in four steps: Plan, Action, Update and Process.

ElementPrompt
TimeWhen does this occur?
SpaceWhere does this occur?
ProtagonistWho is involved?
CausalityWhat triggered this?
IntentionWhat is the purpose?

Each prompt asks for a short phrase or None.

Supervised construction

A GPT teacher reads the context sentence by sentence. For each sentence it writes a memory snippet of the five elements, conditioned on earlier sentences, the memory so far, and the carried-over propositions and protagonists.

Each sentence is concatenated with its snippet into a memory trace. Traces are kept only when the teacher’s final answer matches the gold label. An open-source LLM is then fine-tuned on these traces.

results

Better accuracy with prompts and with fine-tuning

Evaluated on PrOntoQA (5-hop), ProofWriter (depth 5), FOLIO and LogiQA 2.0 (NLI) against Standard, CoT, SymbCoT and Logic-LM.

71.79
average accuracy with GPT-4o-mini, best among methods
+23.3
average FOLIO gain over vanilla fine-tuning
155 vs 269
“uncertain” predictions, SituW vs Logic-LM

Prompt-based setting

MethodPrOntoQAProofWriterFOLIOAvg
GPT-3.5-turbo
Standard47.4040.0045.0950.20
CoT67.8049.1757.3557.78
SymbCoT74.2053.6739.7155.86
Logic-LM58.8051.6650.9853.71
SituW82.2054.6753.4363.43
GPT-4o-mini
Standard58.2039.3361.7653.16
CoT72.6044.1664.2261.46
SymbCoT68.2062.0065.6965.29
Logic-LM77.4037.3362.2558.99
SituW79.0064.3372.0671.79

Accuracy (%). Bold: best; underlined: second best.

Supervised setting, FOLIO

0-shot2-shotFine-tuned
ModelDirectDirectCoTVanillaSituWΔ
Llama-3.1 8B47.7836.9544.3338.9260.10+21.18
Llama-3.1 70B61.5867.9874.3838.4276.35+37.93
Llama-3.3 70B64.0465.0267.0035.5474.88+39.34
Mistral 7B47.2944.3351.2351.2362.56+11.33
Qwen2.5 72B67.4969.9555.6771.1478.00+6.86

Instruction-tuned models. Δ: gain over vanilla fine-tuning.

Fine-tuning on SituW traces beats both prompting and vanilla fine-tuning on every backbone. Vanilla fine-tuning of the 70B Llama models falls below few-shot prompting; SituW traces lift them by 38 to 39 points.

LogiQA, zero-shot

SituW improves both standard and CoT prompting with both GPT models, and raises accuracy in every reasoning category: categorical, conjunction, disjunction, necessity and sufficiency.

ModelSettingVanillaSituWΔ
Human–86.63––
GPT-3.5-turboStandard50.4954.81+4.40
CoT56.2058.30+2.10
GPT-4o-miniStandard57.4161.97+4.56
CoT61.4565.78+4.33
LogiQA accuracy by category. Standard plus SituW gains 8.7 to 14.4 points; CoT plus SituW gains 7.6 to 29.4 points, largest on categorical questions.

Fewer “uncertain” answers

Logic-LM predicts “uncertain” 269 times, more than true or false, and gets 42–43% of true and false questions right. SituW commits to true or false more often, predicts “uncertain” 155 times, and reaches 57–60%. Converting text into logical form seems to drop the contextual cues needed to decide.

Left: true and false accuracy, Logic-LM 42.2 and 43.1, SituW 57.1 and 60.2. Right: predicted label counts; Logic-LM predicts 187 true, 144 false, 269 uncertain; SituW predicts 238 true, 201 false, 155 uncertain.

What each part contributes

Causality alone or intention alone adds almost nothing on top of objective indexing. Together they add 10.43 points: the memory needs both to hold a coherent situation.

ObjectiveCausalityIntentionAcc. (%)
✓54.38
✓✓54.41
✓✓54.93
✓✓✓64.81

Open-source models can build and use the memory on their own. Llama3-70B for both construction and inference gives the best open-source result.

Inference model
ConstructionGemma3-27BLlama3-70BGPT-3.5
Gemma3-4B62.1965.1554.78
Gemma3-27B60.6563.7354.78
Llama3-8B60.4363.3354.69
Llama3-70B61.4564.8554.87

Cost

Averaged over PrOntoQA, ProofWriter and FOLIO, SituW uses about as many tokens as SymbCoT and takes 0.57 s per example, for the best accuracy. Building the memory adds generation steps, which is the main cost of the method.

Tokens
MethodAcc. (%)inoutTime (s)
Direct44.1613213.330.107
CoT58.103181760.482
SymbCoT55.864,6554260.419
SituW63.434,7904250.570

cite

BibTeX

@inproceedings{kim-etal-2026-injecting,
  title     = {Injecting Context via Situation Working Memory for
               Logical Reasoning with {LLM}s},
  author    = {Kim, Jieun and Lim, Seoha and Choi, YoungHae and Cho, Sung-Bae},
  booktitle = {Proceedings of the 64th Annual Meeting of the Association
               for Computational Linguistics (Volume 1: Long Papers)},
  year      = {2026},
  url       = {https://aclanthology.org/2026.acl-long.1499/}
}