Home
Categories
EXPLORE
True Crime
Comedy
Society & Culture
Business
Sports
History
TV & Film
About Us
Contact Us
Copyright
© 2024 PodJoint
00:00 / 00:00
Sign in

or

Don't have an account?
Sign up
Forgot password
https://is1-ssl.mzstatic.com/image/thumb/Podcasts115/v4/61/be/2d/61be2d20-f9b8-85e0-82ef-05ed89759a3d/mza_11872843457524475283.png/600x600bb.jpg
O'Reilly Data Show Podcast
O'Reilly Media
15 episodes
3 weeks ago
The O'Reilly Data Show Podcast explores the opportunities and techniques driving big data, data science, and AI.
Show more...
Business
RSS
All content for O'Reilly Data Show Podcast is the property of O'Reilly Media and is served directly from their servers with no modification, redirects, or rehosting. The podcast is not affiliated with or endorsed by Podjoint in any way.
The O'Reilly Data Show Podcast explores the opportunities and techniques driving big data, data science, and AI.
Show more...
Business
https://www.oreilly.com/radar/wp-content/uploads/sites/3/2019/11/Tools-for-machine-learning-development.jpg
Tools for machine learning development
O'Reilly Data Show Podcast
39 minutes 24 seconds
6 years ago
Tools for machine learning development
In this week’s episode of the Data Show, we’re featuring an interview Data Show host Ben Lorica participated in for the Software Engineering Daily Podcast, where he was interviewed by Jeff Meyerson. Their conversation mainly centered around data engineering, data architecture and infrastructure, and machine learning (ML). Here are a few highlights: Tools for productive collaboration A data catalog, at a high level, basically answers questions around the data that’s available and who is using it so an enterprise can understand access patterns. … The term “data catalog” is generally used when you’ve gotten to the point where you have a team of data scientists and you need a place where they can use libraries in a setting where they can collaborate, and where they can share not only models but maybe even data pipelines and features. The more advanced data science platforms will have automation tools built in. … The ideal scenario is the data science platform is not just for prototyping, but also for pushing things to production. Tools for ML development We have tools for software development, and now we’re beginning to hear about tools for machine learning development—there’s a company here at Strata called Comet.ml, and there’s another startup called Verta.ai. But what has really caught my attention is an open source project from Databricks called MLflow. When it first came out, I thought, ‘Oh, yeah, so we don’t have anything like this. Might have a decent chance of success.’ But I didn’t pay close attention until recently; fast forward to today, there are 80 contributors for 40 companies and 200+ companies using it. What’s good about MLflow is that it has three components and you’re free to pick and choose—you can use one, two, or three. Based on their surveys, the most popular component is the one for tracking and managing machine learning experiments. It’s designed to be useful for individual data scientists, but it’s also designed to be used by teams of data scientists, so they have documented use-cases of MLflow where you have a company managing thousands of models and productions.
O'Reilly Data Show Podcast
The O'Reilly Data Show Podcast explores the opportunities and techniques driving big data, data science, and AI.