The next generation of AI isn't just about processing data; it's about creativity. A new study published on arXiv.org pits OpenAI's GPT-5.2 against Qwen-Max in a head-to-head competition focused on a uniquely challenging task: continuing classic Chinese film scripts. The results, while nuanced, paint a clear picture of GPT-5.2's superior performance in this domain. The implications for AI-assisted screenwriting and creative content generation are significant.
The study, which dropped on January 22nd, 2026, introduces a novel benchmark consisting of 53 iconic Chinese films. Researchers designed an evaluation framework analyzing how well each LLM could generate the second half of a script given only the first. Three samples were generated per film, resulting in a total of 303 valid continuations across both models.
Diving Deep into the Metrics: Structure vs. Raw Similarity
While Qwen-Max showed a slight edge in ROUGE-L scores, a metric measuring lexical similarity, the researchers found GPT-5.2 significantly outperformed in several critical areas. "Qwen-Max achieves marginally higher ROUGE-L (0.2230 vs 0.2114, d=-0.43)," the study notes. However, this advantage was overshadowed by GPT-5.2's dominance in structural preservation, overall quality, and composite scores, leveraging DeepSeek-Reasoner for judging.
GPT-5.2 demonstrated a remarkable ability to maintain the structural integrity of the scripts, earning a score of 0.93 compared to Qwen-Max's 0.75. More impressively, its overall quality score of 44.79 dwarfed Qwen-Max's 25.72, demonstrating a “large effect level” (d>0.8). This success, according to the researchers, boils down to GPT-5.2's strength in preserving character consistency, matching the tone and style of the original script, and adhering to proper formatting conventions.
The Devil is in the Details: Stability and Creative AI
One of the key takeaways from the study is the generation stability, or lack thereof, observed in Qwen-Max. GPT-5.2 consistently delivered high-quality continuations that seamlessly integrated with the existing narrative. In contrast, Qwen-Max exhibited inconsistencies and, at times, struggled to maintain a coherent storyline. This is particularly important in creative applications, where maintaining narrative flow is paramount.
The study authors emphasize the reproducibility of their evaluation framework, making it a valuable tool for future LLM research in Chinese creative writing. This benchmark could become an important standard. It allows other companies to improve their models using an objective way of measuring success. With these developments, we will see AI become a powerful tool for scriptwriters.
"The ability of models like GPT-5.2 to generate high-quality, structurally sound film script continuations suggests a future where AI plays an increasingly prominent role in filmmaking..."
— Dr. Raj Patel, Automatica PressLooking ahead, this research highlights the growing sophistication of large language models and their potential to revolutionize creative industries. While challenges remain, the ability of models like GPT-5.2 to generate high-quality, structurally sound film script continuations suggests a future where AI plays an increasingly prominent role in filmmaking, not as a replacement for human creativity, but as a powerful collaborative tool, augmenting the storytelling process and unlocking new creative possibilities.