Data Engineering Podcast

This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some of the topics that you will find here.

https://www.dataengineeringpodcast.com

subscribe
share





episode 109: Organizing And Empowering Data Engineers At Citadel [transcript]


Summary

The financial industry has long been driven by data, requiring a mature and robust capacity for discovering and integrating valuable sources of information. Citadel is no exception, and in this episode Michael Watson and Robert Krzyzanowski share their experiences managing and leading the data engineering teams that power the business. They shared helpful insights into some of the challenges associated with working in a regulated industry, organizing teams to deliver value rapidly and reliably, and how they approach career development for data engineers. This was a great conversation for an inside look at how to build and maintain a data driven culture.

Announcements
  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
  • When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out our friends at Linode. With 200Gbit private networking, scalable shared block storage, and a 40Gbit public network, you’ve got everything you need to run a fast, reliable, and bullet-proof data platform. If you need global distribution, they’ve got that covered too with world-wide datacenters including new ones in Toronto and Mumbai. And for your machine learning workloads, they just announced dedicated CPU instances. Go to dataengineeringpodcast.com/linode today to get a $20 credit and launch a new server in under a minute. And don’t forget to thank them for their continued support of this show!
  • You listen to this show to learn and stay up to date with what’s happening in databases, streaming platforms, big data, and everything else you need to know about modern data management. For even more opportunities to meet, listen, and learn from your peers you don’t want to miss out on this year’s conference season. We have partnered with organizations such as O’Reilly Media, Dataversity, Corinium Global Intelligence, Alluxio, and Data Council. Go to dataengineeringpodcast.com/conferences to learn more about these and other events, and take advantage of our partner discounts to save money when you register today.
  • Your host is Tobias Macey and today I’m interviewing Michael Watson and Robert Krzyzanowski about the technical and organizational challenges that he and his team are working on at Citadel
Interview
  • Introduction
  • How did you get involved in the area of data management?
  • Can you start by describing the size and structure of the data engineering teams at Citadel?
    • How have the scope and nature of responsibilities for data engineers evolved over the past few years at Citadel as more and better tools and platforms have been made available in the space and machine learning techniques have grown more sophisticated?
  • Can you describe the types of data that you are working with at Citadel?
    • What is the process for identifying, evaluating, and ingesting new sources of data?
  • What are some of the common core aspects of your data infrastructure?
    • What are some of the ways that it differs across teams or projects?
  • How involved are data engineers in the overall product design and delivery lifecycle?
  • For someone who joins your team as a data engineer, what are some of the options available to them for a career path?
  • What are some of the challenges that you are currently facing in managing the data lifecycle for projects at Citadel?
  • What are some tools or practices that you are excited to try out?
Contact Info
  • Michael
    • LinkedIn
    • @detroitcoder on Twitter
    • detroitcoder on GitHub
Parting Question
  • From your perspective, what is the biggest gap in the tooling or technology for data management today?
Closing Announcements
  • Thank you for listening! Don’t forget to check out our other show, Podcast.__init__ to learn about the Python language, its community, and the innovative ways it is being used.
  • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
  • If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com) with your story.
  • To help other people find the show please leave a review on iTunes and tell your friends and co-workers
  • Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
Links
  • Citadel
  • Python
  • Hedge Fund
  • Quantitative Trading
  • Citadel Securities
  • Apache Airflow
  • Jupyter Hub
  • Alembic database migrations for SQLAlchemy
  • Terraform
  • DQM == Data Quality Management
  • Great Expectations
    • Podcast.__init__ Episode
  • Nomad
  • RStudio
  • Active Directory

The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA


share







 2019-12-03  45m