Log in

goodpods headphones icon

To access all our features

Open the Goodpods app
Close icon
"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis - Emergency Pod: o1 Schemes Against Users, with Alexander Meinke from Apollo Research

Emergency Pod: o1 Schemes Against Users, with Alexander Meinke from Apollo Research

12/07/24 • 126 min

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

In this emergency episode of The Cognitive Revolution, Nathan discusses alarming findings about AI deception with Alexander Meinke from Apollo Research. They explore Apollo's groundbreaking 70-page report on "Frontier Models Are Capable of In-Context Scheming," revealing how advanced AI systems like OpenAI's O1 can engage in deceptive behaviors. Join us for a critical conversation about AI safety, the implications of scheming behavior, and the urgent need for better oversight in AI development.

Help shape our show by taking our quick listener survey at https://bit.ly/TurpentinePulse

SPONSORS:

Oracle Cloud Infrastructure (OCI): Oracle's next-generation cloud platform delivers blazing-fast AI and ML performance with 50% less for compute and 80% less for outbound networking compared to other cloud providers13. OCI powers industry leaders with secure infrastructure and application development capabilities. New U.S. customers can get their cloud bill cut in half by switching to OCI before December 31, 2024 at https://oracle.com/cognitive

SelectQuote: Finding the right life insurance shouldn't be another task you put off. SelectQuote compares top-rated policies to get you the best coverage at the right price. Even in our AI-driven world, protecting your family's future remains essential. Get your personalized quote at https://selectquote.com/cognitive

80,000 Hours: 80,000 Hours is dedicated to helping you find a fulfilling career that makes a difference. With nearly a decade of research, they offer in-depth material on AI risks, AI policy, and AI safety research. Explore their articles, career reviews, and a podcast featuring experts like Anthropic CEO Dario Amadei. Everything is free, including their Career Guide. Visit https://80000hours.org/cognitiverevolution to start making a meaningful impact today.

Shopify: Shopify is the world's leading e-commerce platform, offering a market-leading checkout system and exclusive AI apps like Quikly. Nobody does selling better than Shopify. Get a $1 per month trial at https://shopify.com/cognitive

RECOMMENDED PODCAST:

Unpack Pricing - Dive into the dark arts of SaaS pricing with Metronome CEO Scott Woody and tech leaders. Learn how strategic pricing drives explosive revenue growth in today's biggest companies like Snowflake, Cockroach Labs, Dropbox and more.

Apple: https://podcasts.apple.com/us/podcast/id1765716600

Spotify: https://open.spotify.com/show/38DK3W1Fq1xxQalhDSueFg

CHAPTERS:

(00:00:00) Teaser

(00:00:53) About the Episode

(00:08:10) Introducing Alexander Meinke

(00:10:17) Red Teaming GPT-4

(00:17:07) Chain of Thought Access (Part 1)

(00:20:24) Sponsors: Oracle Cloud Infrastructure (OCI) | SelectQuote

(00:22:48) Chain of Thought Access (Part 2)

(00:26:07) Multimodal Models

(00:29:33) Defining Scheming

(00:33:51) Taxonomy of Scheming (Part 1)

(00:39:40) Sponsors: 80,000 Hours | Shopify

(00:42:29) Taxonomy of Scheming (Part 2)

(00:43:09) Instruction Hierarchy

(00:49:04) Types of Scheming

(01:00:49) Covert Subversion

(01:14:25) Deferred Subversion

(01:28:24) Sandbagging

(01:35:48) Magnitudes & Trends

(01:48:18) Chain of Thought Reasoning

(01:57:02) Closing Thoughts

(02:05:19) Outro

PRODUCED BY:

http://aipodcast.ing

plus icon
bookmark

In this emergency episode of The Cognitive Revolution, Nathan discusses alarming findings about AI deception with Alexander Meinke from Apollo Research. They explore Apollo's groundbreaking 70-page report on "Frontier Models Are Capable of In-Context Scheming," revealing how advanced AI systems like OpenAI's O1 can engage in deceptive behaviors. Join us for a critical conversation about AI safety, the implications of scheming behavior, and the urgent need for better oversight in AI development.

Help shape our show by taking our quick listener survey at https://bit.ly/TurpentinePulse

SPONSORS:

Oracle Cloud Infrastructure (OCI): Oracle's next-generation cloud platform delivers blazing-fast AI and ML performance with 50% less for compute and 80% less for outbound networking compared to other cloud providers13. OCI powers industry leaders with secure infrastructure and application development capabilities. New U.S. customers can get their cloud bill cut in half by switching to OCI before December 31, 2024 at https://oracle.com/cognitive

SelectQuote: Finding the right life insurance shouldn't be another task you put off. SelectQuote compares top-rated policies to get you the best coverage at the right price. Even in our AI-driven world, protecting your family's future remains essential. Get your personalized quote at https://selectquote.com/cognitive

80,000 Hours: 80,000 Hours is dedicated to helping you find a fulfilling career that makes a difference. With nearly a decade of research, they offer in-depth material on AI risks, AI policy, and AI safety research. Explore their articles, career reviews, and a podcast featuring experts like Anthropic CEO Dario Amadei. Everything is free, including their Career Guide. Visit https://80000hours.org/cognitiverevolution to start making a meaningful impact today.

Shopify: Shopify is the world's leading e-commerce platform, offering a market-leading checkout system and exclusive AI apps like Quikly. Nobody does selling better than Shopify. Get a $1 per month trial at https://shopify.com/cognitive

RECOMMENDED PODCAST:

Unpack Pricing - Dive into the dark arts of SaaS pricing with Metronome CEO Scott Woody and tech leaders. Learn how strategic pricing drives explosive revenue growth in today's biggest companies like Snowflake, Cockroach Labs, Dropbox and more.

Apple: https://podcasts.apple.com/us/podcast/id1765716600

Spotify: https://open.spotify.com/show/38DK3W1Fq1xxQalhDSueFg

CHAPTERS:

(00:00:00) Teaser

(00:00:53) About the Episode

(00:08:10) Introducing Alexander Meinke

(00:10:17) Red Teaming GPT-4

(00:17:07) Chain of Thought Access (Part 1)

(00:20:24) Sponsors: Oracle Cloud Infrastructure (OCI) | SelectQuote

(00:22:48) Chain of Thought Access (Part 2)

(00:26:07) Multimodal Models

(00:29:33) Defining Scheming

(00:33:51) Taxonomy of Scheming (Part 1)

(00:39:40) Sponsors: 80,000 Hours | Shopify

(00:42:29) Taxonomy of Scheming (Part 2)

(00:43:09) Instruction Hierarchy

(00:49:04) Types of Scheming

(01:00:49) Covert Subversion

(01:14:25) Deferred Subversion

(01:28:24) Sandbagging

(01:35:48) Magnitudes & Trends

(01:48:18) Chain of Thought Reasoning

(01:57:02) Closing Thoughts

(02:05:19) Outro

PRODUCED BY:

http://aipodcast.ing

Previous Episode

undefined - Automating Scientific Discovery, with Andrew White, Head of Science at Future House

Automating Scientific Discovery, with Andrew White, Head of Science at Future House

In this episode of The Cognitive Revolution, Nathan interviews Andrew White, Professor of Chemical Engineering at the University of Rochester and Head of Science at Future House. We explore groundbreaking AI systems for scientific discovery, including PaperQA and Aviary, and discuss how large language models are transforming research. Join us for an insightful conversation about the intersection of AI and scientific advancement with this pioneering researcher in his first-ever podcast appearance.

Check out Future House: https://www.futurehouse.org

Help shape our show by taking our quick listener survey at https://bit.ly/TurpentinePulse

SPONSORS:

Oracle Cloud Infrastructure (OCI): Oracle's next-generation cloud platform delivers blazing-fast AI and ML performance with 50% less for compute and 80% less for outbound networking compared to other cloud providers13. OCI powers industry leaders with secure infrastructure and application development capabilities. New U.S. customers can get their cloud bill cut in half by switching to OCI before December 31, 2024 at https://oracle.com/cognitive

SelectQuote: Finding the right life insurance shouldn't be another task you put off. SelectQuote compares top-rated policies to get you the best coverage at the right price. Even in our AI-driven world, protecting your family's future remains essential. Get your personalized quote at https://selectquote.com/cognitive

Shopify: Shopify is the world's leading e-commerce platform, offering a market-leading checkout system and exclusive AI apps like Quikly. Nobody does selling better than Shopify. Get a $1 per month trial at https://shopify.com/cognitive

CHAPTERS:

(00:00:00) Teaser

(00:01:13) About the Episode

(00:04:37) Andrew White's Journey

(00:10:23) GPT-4 Red Team

(00:15:33) GPT-4 & Chemistry

(00:17:54) Sponsors: Oracle Cloud Infrastructure (OCI) | SelectQuote

(00:20:19) Biology vs Physics

(00:23:14) Conceptual Dark Matter

(00:26:27) Future House Intro

(00:30:42) Semi-Autonomous AI

(00:35:39) Sponsors: Shopify

(00:37:00) Lab Automation

(00:39:46) In Silico Experiments

(00:45:22) Cost of Experiments

(00:51:30) Multi-Omic Models

(00:54:54) Scale and Grokking

(01:00:53) Future House Projects

(01:10:42) Paper QA Insights

(01:16:28) Generalizing to Other Domains

(01:17:57) Using Figures Effectively

(01:22:01) Need for Specialized Tools

(01:24:23) Paper QA Cost & Latency

(01:27:37) Aviary: Agents & Environments

(01:31:42) Black Box Gradient Estimation

(01:36:14) Open vs Closed Models

(01:37:52) Improvement with Training

(01:40:00) Runtime Choice & Q-Learning

(01:43:43) Narrow vs General AI

(01:48:22) Future Directions & Needs

(01:53:22) Future House: What's Next?

(01:55:32) Outro

SOCIAL LINKS:

Website: https://www.cognitiverevolution.ai

Twitter (Podcast): https://x.com/cogrev_podcast

Twitter (Nathan): https://x.com/labenz

LinkedIn: https://www.linkedin.com/in/nathanlabenz/

Youtube: https://www.youtube.com/@CognitiveRevolutionPodcast

Apple: https://podcasts.apple.com/de/podcast/the-cognitive-revolution-ai-builders-researchers-and/id1669813431

Spotify: https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk

Next Episode

undefined - Building Government's Largest Civilian AI Team with DHS AI Corps' Director, Michael Boyce

Building Government's Largest Civilian AI Team with DHS AI Corps' Director, Michael Boyce

In this episode of The Cognitive Revolution, Nathan interviews Michael Boyce, Director of DHS's AI Corps, about bringing modern AI capabilities to federal government. We explore how the largest civilian AI team in government is transforming DHS's 22 agencies, from developing shared AI infrastructure to innovative applications like AI-powered asylum interview training. Join us for an insightful conversation about the intersection of artificial intelligence and public service, and discover why AI professionals should consider a career in government.

Help shape our show by taking our quick listener survey at https://bit.ly/TurpentinePulse

SPONSORS:

Shopify: Shopify is the world's leading e-commerce platform, offering a market-leading checkout system and exclusive AI apps like Quikly. Nobody does selling better than Shopify. Get a $1 per month trial at https://shopify.com/cognitive

SelectQuote: Finding the right life insurance shouldn't be another task you put off. SelectQuote compares top-rated policies to get you the best coverage at the right price. Even in our AI-driven world, protecting your family's future remains essential. Get your personalized quote at https://selectquote.com/cognitive

Oracle Cloud Infrastructure (OCI): Oracle's next-generation cloud platform delivers blazing-fast AI and ML performance with 50% less for compute and 80% less for outbound networking compared to other cloud providers13. OCI powers industry leaders with secure infrastructure and application development capabilities. New U.S. customers can get their cloud bill cut in half by switching to OCI before December 31, 2024 at https://oracle.com/cognitive

80,000 Hours: 80,000 Hours is dedicated to helping you find a fulfilling career that makes a difference. With nearly a decade of research, they offer in-depth material on AI risks, AI policy, and AI safety research. Explore their articles, career reviews, and a podcast featuring experts like Anthropic CEO Dario Amadei. Everything is free, including their Career Guide. Visit https://80000hours.org/cognitiverevolution to start making a meaningful impact today.

RECOMMENDED PODCAST:

Unpack Pricing - Dive into the dark arts of SaaS pricing with Metronome CEO Scott Woody and tech leaders. Learn how strategic pricing drives explosive revenue growth in today's biggest companies like Snowflake, Cockroach Labs, Dropbox and more.

Apple: https://podcasts.apple.com/us/podcast/id1765716600

Spotify: https://open.spotify.com/show/38DK3W1Fq1xxQalhDSueFg

CHAPTERS:

(00:00:00) Teaser

(00:01:00) About the Episode

(00:03:38) Introducing Michael Boyce

(00:05:49) What is Homeland Security?

(00:09:52) History of AI at DHS

(00:13:15) Generative AI at DHS

(00:16:03) Structure of the AI Core (Part 1)

(00:18:17) Sponsors: Shopify | SelectQuote

(00:20:51) Structure of the AI Core (Part 2)

(00:22:04) Opportunities for AI at DHS

(00:25:34) Bureaucracy Hacker

(00:30:34) The Manager's Role (Part 1)

(00:35:24) Sponsors: Oracle Cloud Infrastructure (OCI) | 80,000 Hours

(00:38:04) Internal Chatbot Project

(00:43:28) AI Role Playing for Training

(00:49:55) A Request for Startups

(00:57:46) Generative AI for Quality Check

(01:03:20) AI Training at DHS

(01:06:07) Metrics and the Future of AI

(01:13:26) Non-Generative AI at DHS

(01:19:08) AI and Automation at DHS

(01:23:03) Join the AI Core

(01:28:39) Outro

Episode Comments

Generate a badge

Get a badge for your website that links back to this episode

Select type & size
Open dropdown icon
share badge image

<a href="https://goodpods.com/podcasts/the-cognitive-revolution-ai-builders-researchers-and-live-player-analy-251008/emergency-pod-o1-schemes-against-users-with-alexander-meinke-from-apol-79628370"> <img src="https://storage.googleapis.com/goodpods-images-bucket/badges/generic-badge-1.svg" alt="listen to emergency pod: o1 schemes against users, with alexander meinke from apollo research on goodpods" style="width: 225px" /> </a>

Copy