
#TransformerArchitecture #AttentionMechanism #LLMs Encoders, cross attention and masking for LLMs: SuperDataScience Founder Kirill Eremenko returns to the SuperDataScience podcast, where he speaks with @JonKrohnLearns about transformer architectures and why they are a new frontier for generative AI. If you’re interested in applying LLMs to your business portfolio, you’ll want to pay close attention to this episode! This episode is brought to you by Ready Tensor, where innovation meets reproducibility (https://www.readytensor.ai/), by Oracle NetSuite business software (https://netsuite.com/superdata), and by Intel and HPE Ezmeral Software Solutions (https://hpe.com/ezmeral/chatbots). Interested in sponsoring a SuperDataScience Podcast episode? Visit https://passionfroot.me/superdatascience for sponsorship information. In this episode you will learn: • [00:00:00] Introduction • [00:13:04] How decoder-only transformers work • [00:38:07] How cross-attention works in transformers • [00:50:00] How encoders and decoders work together (an example) • [01:17:52] How encoder-only architectures excel at understanding natural language • [01:24:16] The importance of masking during self-attention Additional materials: https://www.superdatascience.com/759

The AI Model That Finally Beat Me (with Jon Krohn)

Ten Years of the Super Data Science Podcast (Ep. 1000 with Jon, Kirill and Special Guests)

917: 8 Steps to Becoming an AI Engineer — with Kirill Eremenko

899: Landing $200k+ AI Roles: Real Cases from the SuperDataScience Community — with Kirill Eremenko

853: Generative AI for Business — with Kirill Eremenko and Hadelin de Ponteves