Artwork

Inhalt bereitgestellt von Changelog Media. Alle Podcast-Inhalte, einschließlich Episoden, Grafiken und Podcast-Beschreibungen, werden direkt von Changelog Media oder seinem Podcast-Plattformpartner hochgeladen und bereitgestellt. Wenn Sie glauben, dass jemand Ihr urheberrechtlich geschütztes Werk ohne Ihre Erlaubnis nutzt, können Sie dem hier beschriebenen Verfahren folgen https://de.player.fm/legal.
Player FM - Podcast-App
Gehen Sie mit der App Player FM offline!

Towards high-quality (maybe synthetic) datasets

57:04
 
Teilen
 

Manage episode 444356710 series 2385063
Inhalt bereitgestellt von Changelog Media. Alle Podcast-Inhalte, einschließlich Episoden, Grafiken und Podcast-Beschreibungen, werden direkt von Changelog Media oder seinem Podcast-Plattformpartner hochgeladen und bereitgestellt. Wenn Sie glauben, dass jemand Ihr urheberrechtlich geschütztes Werk ohne Ihre Erlaubnis nutzt, können Sie dem hier beschriebenen Verfahren folgen https://de.player.fm/legal.

As Argilla puts it: “Data quality is what makes or breaks AI.” However, what exactly does this mean and how can AI team probably collaborate with domain experts towards improved data quality? David Berenstein & Ben Burtenshaw, who are building Argilla & Distilabel at Hugging Face, join us to dig into these topics along with synthetic data generation & AI-generated labeling / feedback.

Join the discussion

Changelog++ members save 11 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • Fly.ioThe home of Changelog.com — Deploy your apps close to your users — global Anycast load-balancing, zero-configuration private networking, hardware isolation, and instant WireGuard VPN connections. Push-button deployments that scale to thousands of instances. Check out the speedrun to get started in minutes.
  • WorkOSA platform that gives developers a set of building blocks for quickly adding enterprise-ready features to their application. Add Single Sign-On (Okta, Azure, Google, Microsoft OAuth), sync users from any SCIM directory, HRIS integration, audit trails (SIEM), free magic link sign-in. WorkOS is designed for developers and offers a single, elegant interface that abstracts dozens of enterprise integrations. Learn more and get started at WorkOS.com
  • Eight SleepTake your sleep and recovery to the next level. Go to eightsleep.com/PRACTICALAI and use the code PRACTICALAI to get $350 off your very own Pod 4 Ultra. You can try it for free for 30 days - but we’re confident you will not want to return it. Once you experience AI-optimized sleep, you’ll wonder how you ever slept without it. Currently shipping to: United States, Canada, United Kingdom, Europe, and Australia.

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

  continue reading

Kapitel

1. Welcome to Practical AI (00:00:00)

2. Sponsor: Fly (00:00:44)

3. What does data collaboration mean? (00:03:56)

4. Understanding your data (00:07:18)

5. How to start curating data (00:09:58)

6. Practical steps to scale (00:13:12)

7. Sponsor: WorkOS (00:16:52)

8. Traditional & new usecases (00:20:23)

9. Virtues of smaller models (00:24:51)

10. What Argilla looks like (00:27:04)

11. User backgrounds (00:30:55)

12. The non-technical POV (00:34:21)

13. Sponsor: Eight Sleep (00:38:23)

14. AI feedback (00:41:09)

15. Hallucination issues (00:44:50)

16. What is Distilabel (00:46:10)

17. Usage & adoption (00:50:08)

18. Where things are going (00:52:55)

19. This is muy bueno (00:55:34)

20. Outro (00:56:15)

292 Episoden

Artwork
iconTeilen
 
Manage episode 444356710 series 2385063
Inhalt bereitgestellt von Changelog Media. Alle Podcast-Inhalte, einschließlich Episoden, Grafiken und Podcast-Beschreibungen, werden direkt von Changelog Media oder seinem Podcast-Plattformpartner hochgeladen und bereitgestellt. Wenn Sie glauben, dass jemand Ihr urheberrechtlich geschütztes Werk ohne Ihre Erlaubnis nutzt, können Sie dem hier beschriebenen Verfahren folgen https://de.player.fm/legal.

As Argilla puts it: “Data quality is what makes or breaks AI.” However, what exactly does this mean and how can AI team probably collaborate with domain experts towards improved data quality? David Berenstein & Ben Burtenshaw, who are building Argilla & Distilabel at Hugging Face, join us to dig into these topics along with synthetic data generation & AI-generated labeling / feedback.

Join the discussion

Changelog++ members save 11 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • Fly.ioThe home of Changelog.com — Deploy your apps close to your users — global Anycast load-balancing, zero-configuration private networking, hardware isolation, and instant WireGuard VPN connections. Push-button deployments that scale to thousands of instances. Check out the speedrun to get started in minutes.
  • WorkOSA platform that gives developers a set of building blocks for quickly adding enterprise-ready features to their application. Add Single Sign-On (Okta, Azure, Google, Microsoft OAuth), sync users from any SCIM directory, HRIS integration, audit trails (SIEM), free magic link sign-in. WorkOS is designed for developers and offers a single, elegant interface that abstracts dozens of enterprise integrations. Learn more and get started at WorkOS.com
  • Eight SleepTake your sleep and recovery to the next level. Go to eightsleep.com/PRACTICALAI and use the code PRACTICALAI to get $350 off your very own Pod 4 Ultra. You can try it for free for 30 days - but we’re confident you will not want to return it. Once you experience AI-optimized sleep, you’ll wonder how you ever slept without it. Currently shipping to: United States, Canada, United Kingdom, Europe, and Australia.

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

  continue reading

Kapitel

1. Welcome to Practical AI (00:00:00)

2. Sponsor: Fly (00:00:44)

3. What does data collaboration mean? (00:03:56)

4. Understanding your data (00:07:18)

5. How to start curating data (00:09:58)

6. Practical steps to scale (00:13:12)

7. Sponsor: WorkOS (00:16:52)

8. Traditional & new usecases (00:20:23)

9. Virtues of smaller models (00:24:51)

10. What Argilla looks like (00:27:04)

11. User backgrounds (00:30:55)

12. The non-technical POV (00:34:21)

13. Sponsor: Eight Sleep (00:38:23)

14. AI feedback (00:41:09)

15. Hallucination issues (00:44:50)

16. What is Distilabel (00:46:10)

17. Usage & adoption (00:50:08)

18. Where things are going (00:52:55)

19. This is muy bueno (00:55:34)

20. Outro (00:56:15)

292 Episoden

Alle Folgen

×
 
Loading …

Willkommen auf Player FM!

Player FM scannt gerade das Web nach Podcasts mit hoher Qualität, die du genießen kannst. Es ist die beste Podcast-App und funktioniert auf Android, iPhone und im Web. Melde dich an, um Abos geräteübergreifend zu synchronisieren.

 

Kurzanleitung