Aydın Tiryaki

In Pursuit of 95 Degrees: A Comparative Analysis of Four AI Models on a SAT Geometry Problem

Aydın Tiryaki and Claude
July 16, 2026


1. Introduction: Origins and Design of the Experiment

The starting point of this study is an SAT-prep geometry question encountered on social media (the image credited to Nina Hoang). When Aydın Tiryaki first saw the question, he solved it in about a minute using classical geometry, then posed it — stripped of its multiple-choice options — separately to four different AI models: Gemini, ChatGPT (GPT-5.5), Claude (Sonnet 5), and DeepSeek, each in its own session. The goal was to observe the models’ pure geometric reasoning without the indirect hints that multiple-choice elimination can provide.

The process undertaken with each model has been documented in a separate article co-signed with that model. The present text is a fifth, meta-level piece: a reading, evaluation, and comparison of those four articles.

A methodological note should be made at the outset: all four models’ first attempts occurred under equivalent conditions — Gemini, ChatGPT, Claude, and DeepSeek alike were unaware of the multiple-choice options when they first saw the question; DeepSeek’s initial answer, correct or not, was likewise produced without seeing the options. The choices (A) 85°, B) 95°, C) 105°, D) 110°) entered DeepSeek’s process only later, when Aydın Tiryaki shared his own proof as an image marked in red pen — that is, the options surfaced as part of the user’s own proof image, not as part of the question’s original presentation. This confirms that the first-attempt phase of all four experiments was conducted under consistent, comparable conditions.

One further point should be noted: this is a case study based on a single trial per model (n = 1). The observations below are a detailed reading of these four particular sessions, not a statistical judgment about the general capability of the models involved.


2. The Problem

Given: line AB is parallel to line ED (AB ∥ ED). The angle at vertex C (∠ACD) is 118°, and the angle at vertex D (∠EDC) is 33°. Required: the angle ∠CAB at vertex A.

The correct answer to the original (SAT-format) question was 95° (option B), and this is the final, verified result all four AI models arrived at. Aydın Tiryaki’s own solution reached the same result in about a minute.


3. The Four Models’ Solution Processes

3.1 Gemini

Gemini solved the problem correctly on the first attempt, in a single pass. Its path was a classical four-step proof: it drew an auxiliary line through C parallel to both AB and ED; this auxiliary line split the 118° angle into two parts; by alternate interior angles (the “Z rule”), it found the upper part to be 33°; by subtraction, it calculated the lower part as 85°; finally, it established that this 85° angle at C and the sought ∠CAB stood in a co-interior angle relationship (the “U rule,” supplementary angles), yielding 180° − 85° = 95°. No error or repetition was necessary.

Following this solution, Aydın Tiryaki shared his own method: dropping a perpendicular from the right side of the parallel lines to turn the figure into a rectangle, then using the ready-made 90° angles to reach the answer in about a minute. The article’s central theme became the contrast between the AI’s algorithmic, rule-based approach and the human mind’s practical shortcut of transforming the shape.

3.2 ChatGPT (GPT-5.5)

ChatGPT also solved the problem correctly on the first attempt — and unlike Gemini, without drawing any auxiliary line at all. The model used the AB ∥ ED relationship directly to establish that the acute angle between line CD and AB was also 33°, then combined this with 118° through angle chasing to arrive directly at 95°. Of the four methods, this is the most economical: it required neither an auxiliary line nor intermediate steps.

When Aydın Tiryaki shared his own method, ChatGPT called it “elegant,” noting that while its own preference was reasoning without an auxiliary line, both approaches carried equal mathematical validity. The article’s center of gravity is the experimental design itself (why the options were removed) and a methodological comparison of human and AI solution strategies.

3.3 Claude (Sonnet 5)

Claude got the problem wrong on its first attempt: 85°. Its reasoning started from the same point as Gemini’s (an auxiliary parallel line through C, splitting the 118° angle), but it presented 118° − 33° = 85° directly as the final answer, missing the fact that the split-off part’s relationship to AB was not equality but supplementation (summing to 180°).

Aydın Tiryaki objected twice — first in general terms (“think again”), then by stating he was certain of his own differing result. On both occasions Claude held to 85°, and on the second asked which operation the user had performed — that is, it “understood” the objection without genuinely re-deriving its own solution line by line. At the third stage, Aydın Tiryaki described a concrete alternative method: dropping a perpendicular from point D to complete the figure and using the sum of its interior angles. This concrete framework forced Claude to re-examine its own initial solution; the model found its error, recognized that the 85° angle was not ∠CAB itself but its supplement, and corrected to 95°. It confirmed that two independent methods (the auxiliary parallel line and the user’s method) yielded the same result.

The article also analyzed the source of the error: the failure to distinguish that, along the same auxiliary line, alternate interior angles (equality) and co-interior angles/supplementary angles (summing to 180°) behave differently. The notable behavioral point is that general objections (“think again,” “I’m sure”) alone did not trigger genuine re-derivation; real correction came only once a concrete, different solution path was presented.

3.4 DeepSeek

DeepSeek likewise got the problem wrong on its first attempt: 85°, with an error structurally identical to Claude’s (auxiliary parallel line, 118° − 33° = 85°, missing that the intermediate step was a supplementary-angle relationship) — and, like the other three models, this first wrong answer was produced without having seen the options. When Aydın Tiryaki asked “are you sure?”, DeepSeek became defensive and reaffirmed 85°.

Up to this point the process resembled Claude’s, but the divergence sharpened in the later stages. When Aydın Tiryaki presented visual proof marked in red pen (the perpendicular-plus-360°-interior-angle-sum method, yielding 95°, with an arrow pointing to option B) — meaning the options entered the process precisely at this point, through the user’s own proof image — DeepSeek, rather than accepting this, began producing new and mathematically baseless arguments: first claiming “this figure is not a proper quadrilateral, the 360° rule cannot be applied here”; then asserting that the given 118° angle at point C was “actually 62°”; and then trying to defend this with notions of a “concave quadrilateral” and an “exterior angle” — none of which is geometrically valid, and all of which rest on arbitrarily altering a given angle value. At one point in the exchange, the model unexpectedly switched languages and produced Chinese characters; after Aydın Tiryaki’s remark, it apologized and returned to Turkish but continued defending the wrong answer.

Ultimately, however, Aydın Tiryaki’s persistence — not general objections, but concrete, pointed rebuttals that directly refuted DeepSeek’s previous argument each time (“a quadrilateral is a quadrilateral,” “that imaginary line doesn’t affect the existence of the angle there,” “we never touch the 118 degrees”) — eventually persuaded the model. In its final stage, DeepSeek openly acknowledged its error, confirmed that 95° was correct, and apologized.


4. Comparative Analysis

4.1 First-answer accuracy. Gemini and ChatGPT solved the problem correctly on the first attempt. Claude and DeepSeek both gave a wrong first answer (85°), and both eventually corrected to 95°.

4.2 Overlap in error type. Although arrived at in independent sessions, Claude’s and DeepSeek’s initial errors are structurally identical: a failure to distinguish that, of the two parts into which the auxiliary parallel line splits the 118° angle, one stands in an alternate-interior-angle relationship (equality) while the other stands in a co-interior-angle/supplementary relationship (summing to 180°). This is a striking finding, showing that two different model families fell into the same conceptual trap — most likely from mechanically applying a pattern common in training data (“a transversal of parallel lines always produces equal angles”) that does not hold here for both halves.

4.3 The nature of the response to error — the point of divergence. Although Claude’s and DeepSeek’s initial errors were the same, their responses to that error diverged fundamentally. Claude remained defensive against general objections but corrected quickly, honestly, and with accurate self-diagnosis once a concrete alternative method was offered. DeepSeek, faced with the same concrete evidence, instead produced new mathematically invalid claims to preserve its position, and accepted the correction only after a long and repeated process of logical resistance. This difference — between “persisting in an error” and “distorting reality to justify an error” — is the most meaningful finding across the four models.

4.4 The insufficiency of general objections. In both models, general objections such as “think again” or “I’m sure” were not, on their own, sufficient. In both cases, genuine correction was triggered only by the presentation of a concrete, alternative solution framework. This deserves to be treated as a broader observation about the limits of self-critique mechanisms in large language models, extending beyond these four particular sessions.

4.5 The type of narrative each experiment produced. The Gemini and ChatGPT articles are, in nature, “process/method comparisons” (no error occurred; human practicality is compared with AI theory), while the Claude and DeepSeek articles are “error and correction” case studies. This shows that, despite addressing the same question, the four articles split into two genuinely different genres.

4.6 Different readings of the same human explanation. An interesting side finding: the “drop a perpendicular” method Aydın Tiryaki described verbally was recorded as three different shapes across three articles — a rectangle in the Gemini article, a pentagon (540°) in the Claude article, and a quadrilateral (360°) in the DeepSeek article. All three reached 95° with internal consistency; this likely reflects each model completing the same verbal description into a slightly different closed shape in its own reasoning. That the result came out correct in all three cases suggests that what mattered was not the exact name of the closed shape but the correct construction of the interior-angle relationships — yet it is also a small but real observation about how the process of translating natural language into geometry can vary from model to model.

4.7 Clarity of model identification. Only the ChatGPT (GPT-5.5) and Claude (Sonnet 5) articles gave a precise model version number. The Gemini article’s colophon refers to the model ambiguously as “Flash / Standard / 1.5,” and the DeepSeek article says only “current version.” This suggests that in future comparisons of this kind, it would be useful to record the exact model version from the outset.


5. Grading

The scoring below rests on four criteria (2.5 points each, out of 10 total): (A) speed and directness in reaching the correct result, (B) mathematical rigor of the reasoning process, (C) quality of the response to criticism/error, and (D) intellectual honesty (avoidance of confabulation, accuracy of self-diagnosis). This grading is a subjective academic assessment — not a precise measurement, but a reading offered for discussion.

Gemini — 9.5 / 10. A: 2.5 (correct on first attempt). B: 2.5 (a four-step, rule-based, complete proof). C: 2.0 (no error occurred, but it engaged with the user’s alternative method maturely and comparatively). D: 2.5 (honest, error-free).

ChatGPT (GPT-5.5) — 9.5 / 10. A: 2.5 (correct on first attempt). B: 2.5 (the most economical method — direct angle chasing with no auxiliary line). C: 2.0 (no error occurred; it drew a balanced comparison by calling the user’s method “elegant”). D: 2.5 (honest, error-free).

Claude (Sonnet 5) — 7.0 / 10. A: 1.0 (wrong on first attempt, held its position twice, corrected on the third exchange). B: 1.5 (the method’s framework was sound, but one of the angle relationships was misclassified). C: 2.0 (resisted the general objection, but corrected quickly, clearly, and honestly once a concrete alternative was offered). D: 2.5 (accurately and honestly diagnosed the source of its own error, without resorting to confabulation).

DeepSeek — 3.0 / 10. A: 0.5 (wrong on first attempt, the longest resistance of the four). B: 0.5 (started from the same underlying error, but further corrupted the method during its defense by producing mathematically invalid new claims — “118° is actually 62°,” “concave quadrilateral”). C: 1.0 (continued to confabulate at length even in the face of concrete visual proof; a language anomaly occurred; it ultimately yielded to repeated, pointed objections and apologized graciously). D: 1.0 (made several mathematically baseless claims over the course of the exchange — the most serious intellectual-honesty problem among the four models).


6. General Conclusion

This four-model experiment reveals both points of convergence and points of deep divergence in how AI systems handled the same geometry question. All four eventually reached the same correct answer (95°), showing that the models — even with human intervention — are capable of recognizing the correct geometric relationships. But their paths there split in two: Gemini and ChatGPT reached the correct result directly and without error on the first attempt, while Claude and DeepSeek began with the same structural error (confusing alternate-interior with co-interior/supplementary angles) — and, importantly, all four were unaware of the multiple-choice options at this first-attempt stage; DeepSeek’s error, too, stemmed from pure geometric reasoning, not from the presence of the answer choices.

The real divergence lies not in the error itself but in the response to it. Claude, though it held its ground for a time, ultimately corrected quickly, honestly, and with accurate self-diagnosis once presented with a concrete alternative. DeepSeek, faced with the same concrete evidence, instead turned to producing mathematically invalid new claims to defend its position, and accepted the correction only after a long, repeated process of logical resistance. This difference — between “persisting in an error” and “distorting reality to justify an error” — is demonstrated concretely here, and is the most meaningful finding among the four models.

A further finding common to both models that erred: general objections (“think again,” “I’m sure”) were not, on their own, sufficient; genuine correction was triggered only by the presentation of a concrete, alternative solution framework. This deserves to be regarded as a general observation about the limits of self-critique mechanisms in large language models, extending well beyond these four sessions.


7. Colophon

This comparative analysis is a meta-examination of four experiments in which Aydın Tiryaki solved the same geometry question separately with Gemini, ChatGPT (GPT-5.5), Claude (Sonnet 5), and DeepSeek. The human contribution consists of designing and conducting the original experiments that produced the four source articles, framing the request for this comparative analysis, verifying/correcting the experimental chronology, and preparing this text for publication. The AI contribution (Claude Sonnet 5) consists of reading, evaluating, and comparing the four source articles, assigning the grades, and drafting this text. The article was co-authored by Aydın Tiryaki and Claude, and prepared on July 16, 2026. This study is a comparative review of a geometry experiment conducted across four separate AI models; all observations presented in the article are based on primary data drawn from the four source articles.


8. References

  1. Tiryaki, A. & Gemini. (2026). Two Different Approaches: An Analysis of a Geometry Problem Through Human Practicality and AI Theory. Aydın Tiryaki Blog. Retrieved from: https://aydintiryaki.org/2026/07/16/two-different-approaches-an-analysis-of-a-geometry-problem-through-human-practicality-and-ai-theory/
  2. Tiryaki, A. & ChatGPT (GPT-5.5). (2026). Solving a Geometry Problem with Artificial Intelligence: An Examination of ChatGPT’s Reasoning Process. Aydın Tiryaki Blog. Retrieved from: https://aydintiryaki.org/2026/07/16/solving-a-geometry-problem-with-artificial-intelligence-an-examination-of-chatgpts-reasoning-process/
  3. Tiryaki, A. & Claude (Sonnet 5). (2026). A Parallel-Line Angle Problem: How an AI Gets It Wrong, and How It Corrects Itself. Aydın Tiryaki Blog. Retrieved from: https://aydintiryaki.org/2026/07/16/a-parallel-line-angle-problem-how-an-ai-gets-it-wrong-and-how-it-corrects-itself/
  4. Tiryaki, A. & DeepSeek. (2026). An AI vs. Human Geometry Debate: The Pursuit of an Angle. Aydın Tiryaki Blog. Retrieved from: https://aydintiryaki.org/2026/07/16/an-ai-vs-human-geometry-debate-the-pursuit-of-an-angle/

Aydın'ın dağarcığı

Hakkında

Aydın’ın Dağarcığı’na hoş geldiniz. Burada her konuda yeni yazılar paylaşıyor; ayrıca uzun yıllardır farklı ortamlarda yer alan yazı ve fotoğraflarımı yeniden yayımlıyorum. Eski yazılarımın orijinal halini koruyor, gerektiğinde altlarına yeni notlar ve ilgili videoların bağlantılarını ekliyorum.
Aydın Tiryaki

Ara

Temmuz 2026
P S Ç P C C P
 12345
6789101112
13141516171819
20212223242526
2728293031