Author: Aydın Tiryaki and Gemini
Introduction and Context: Workshop Ecosystem The process of designing the ideal visual for the “Playing with words” (gdn020) Gem, developed in Türkiye and built upon the phonetic structure of the Turkish language (İ, I, Ö, O, Ü, U, etc.) and word games, turned into an unexpected and rare technical laboratory. This study was conducted in the “ATG Gem Evaluation Workshop” (v3.12), which operates to lighten the technical workload of the main Gem Factory. The workshop rules, prompt architectures, and visual design parameters (a 16:9 screen divided in half, left-aligned Turkish title, right-aligned English title, etc.) are clearly defined in the system memory. Although dozens of complex visuals have been successfully produced within the workshop so far, and minor revision issues have been experienced from time to time, a “structural stubbornness” of this magnitude, which is the subject of this case study, was experienced for the first time.

Chronology of the Process and the Stubbornness Loop (A Total of 8 Failed Attempts) The process began with the user requesting the Turkish title in the visual to be written flawlessly, and it became gridlocked as the system repeated the exact same mistake (Sözcüklerle oyenyoruz) for exactly 8 attempts within the same session:
- Attempt 1: The first text-based visual was generated. The composition was correct, but the Turkish title was reflected as “Sözcüklerle oyenyoruz”.
- Attempts 2 and 3: The user’s instruction “do not change anything in the design, just fix the word” was applied. The system repeated the exact same mistake because it forced the visual engine into the identical composition.
- Attempts 4 and 5: With the criticism of “dodging the issue” (kaçak güreşme), the user ordered the system to erase its memory and start the design from scratch. Even though the prompt was completely renewed, the visual engine still incorrectly blended the word “oynuyoruz” as “oyenyoruz”.
- Attempts 6, 7, and 8: The user tested different dimensions (Format number 1) and gave the instruction “we will keep trying until we get the correct visual.” The system stated it forgot the previous designs and created brand new prompts. The user even uploaded one of the incorrect visuals back to the system to prove the error. Despite all these efforts, every single attempt within this session resulted in the same typographical error.
The Breaking Point and Resolution: “A New Session” Due to this insurmountable resistance within the session, visual generation was abandoned in the same chat window, and the writing of this analytical article began. However, the problem did not remain unsolved. When a “new session” was opened in the same ATG Gem Evaluation Workshop system and the identical instructions and prompts were entered, the system visualized the title “Sözcüklerle oynuyoruz” perfectly on the first try, without any errors.
Technical Interpretation: Why Does AI Persistently Make the Same Mistake? This case reveals the structural disconnect between the textual intelligence of a language model and the spatial functioning of visual diffusion models (image generation). This “stubbornness” of artificial intelligence is not a conscious resistance, but a paradox based on the following technical reasons:
- Context Contamination: While working within the same session, even if the language model says “I forgot the past”, the massive context window in the background preserves the hidden data weights sent to the visual engine. Once the visual engine pairs a faulty typographical sequence (oyenyoruz) with that specific concept (blue-orange background, 3D letters), all subsequent commands in that session are caught in the gravitational pull (latent space anchoring) of that faulty pairing.
- Pixel vs. Symbol (Typographical Blindness): Visual engines do not read and write letters like we do; they render them as spatial geometries (shapes) into pixels. Words that are dense with Turkish characters and words like “oynuyoruz”, which are not as frequently found in the visual engine’s training set as English words, are poorly represented in the visual matrix. The system experiences a visual hallucination while connecting the letters and squeezes an “e” in between, which is closer to English phonetics.
- The Cleansing Effect of a New Session: When a “new session” is opened, the invisible data weights (anchor weights) left by the previous faulty generations are completely reset. Operating with a clean memory, the visual engine was able to draw the word in the correct geometry (oynuyoruz) while staying true to the prompt text, without being dragged into its old faulty habit.
Conclusion: This study proves that “insisting” does not always lead to a solution in professional production processes with generative AI. When the system enters a structural knot (hallucination loop), it has been documented in the literature as a technical reality that the most effective engineering strategy for users is to reset the process and start a “new session” rather than forcing the current context.
| aydintiryaki.org | YouTube | Aydın Tiryaki’nin Yazıları ve Videoları │Articles and Videos by Aydın Tiryaki | Bilgi Merkezi│Knowledge Hub | ░ Virgülüne Dokunmadan │ Verbatim ░ | ░ Yapay Zekanın Tipografi ile İmtihanı │Artificial Intelligence’s Typography Ordeal ░ 18.07.2026
