LinkedIn, a long-time leader in AI-driven recommender systems, has revealed a surprising pivot in its approach to next-generation AI. Instead of relying on prompting large language models (LLMs), the company found superior results by focusing on smaller, highly specialized models fine-tuned with a multi-teacher distillation technique. This strategy, according to Erran Berger, VP of product engineering at LinkedIn, proved to be a "breakthrough" in achieving the desired accuracy, latency, and efficiency for matching job seekers with opportunities.

Why Prompting Fell Short

LinkedIn's initial exploration of LLMs for job recommendations quickly revealed the limitations of prompt engineering. "There was just no way we were gonna be able to do that through prompting," Berger stated on the Beyond the Pilot podcast. The complexity of interpreting job queries, candidate profiles, and job descriptions in real-time demanded a more nuanced and controlled approach than simply feeding prompts into off-the-shelf models could provide. The need to align AI behavior with specific product policies further solidified the decision to move away from general-purpose LLMs.

The company instead turned to a strategy of creating a detailed product policy document, spanning 20-30 pages, to meticulously score job description and profile pairs across multiple dimensions. This document, refined through numerous iterations with the product management team, became the foundation for training a 7-billion-parameter teacher model. That model was then distilled into smaller, more efficient student models optimized for specific tasks.

The Power of Multi-Teacher Distillation

LinkedIn's innovative use of multi-teacher distillation proved to be a key element in their success. The initial product policy-focused teacher model was joined by a second teacher model oriented toward click prediction. This combination allowed them to distill a 1.7 billion parameter model for training purposes. According to Berger, this technique allowed the team to achieve affinity to the original product policy and improve click prediction.

Berger explained the concept with the analogy of a chat agent being trained by two teachers: one focused on accuracy and the other on tone. "By now mixing them, you get better outcomes, but also iterate on them independently," he said. "That was a breakthrough for us."

This modularized and componentized training process allowed for independent iteration and optimization of different aspects of the model's behavior. The result was a system that not only understood the nuances of LinkedIn's product policy but also effectively predicted user engagement.

A New Blueprint for AI Development

Beyond the technical achievements, LinkedIn's experience has also transformed how its teams collaborate on AI projects. Product managers, previously focused on strategy and user experience, now work closely with machine learning engineers to define and refine product policies. This collaborative approach ensures that the AI models are aligned with both business goals and user needs. According to Berger, this new way of working has become "a blueprint for basically any AI products we do at LinkedIn."

""How product managers work with machine learning engineers now is very different from anything we've done previously. It’s now a blueprint for basically any AI products we do at LinkedIn.""

— Erran Berger, VP of product engineering at LinkedIn

LinkedIn's journey highlights the importance of tailoring AI solutions to specific needs, even if it means moving away from the current hype around large language models. By prioritizing control, interpretability, and alignment with product policies, the company has achieved significant improvements in its recommender systems. This approach also underscores the value of close collaboration between product and engineering teams in the development of effective and responsible AI.