Entrepreneurial Geekiness

Ian is a London-based independent Chief Data Scientist who coaches teams, teaches and creates data products. More about Ian here.

Coaching

Training

Jobs

Products

Consulting

Weekish notes

Data science, Python July 5, 2020

I gave another iteration of my Making Pandas Fly talk sequence for PyDataAmsterdam recently and received some lovely postcards from attendees as a result. I’ve also had time to list new iterations of my training courses for Higher Performance Python (October) and Software Engineering for Data Scientists (September), both will run virtually via Zoom & Slack in the UK timezone.

I’ve been using my dtype_diet tool to time more performance improvements with Pandas and I look forward to talking more on this at EuroPython this month.

In baking news I’ve improved my face-making on sourdough loaves (but still have work to do) and I figure now is a good time to have a crack at dried-yeast baking again.

Ian is a Chief Interim Data Scientist via his Mor Consulting. Sign-up for Data Science tutorials in London and to hear about his data science thoughts and jobs. He lives in London, is walked by his high energy Springer Spaniel and is a consumer of fine coffees.

“Making Pandas Fly” for PyDataAmsterdam 2020

Data science, High Performance Python Book, pydata June 23, 2020

I thank the PyDataAmsterdam 2020 organisers for another chance to speak on Making Pandas Fly (PyDataAmsterdam 2020). This variant of the talk focuses more on:

Understanding when categories beat strings and smaller floats beat larger ones
What’s happening with NumPy behind the scenes
How we can save 50% of our RAM (and so fit in more data to the same machine) by checking dtypes with my dtype_diet tool
Considering that float16 is simulated on modern hardware and so is memory efficient but slow for calculating!
Tips to install bottleneck & numexpr to make Pandas faster
Digging into some Pandas internals when I filed a bug – and what I learned as a result (you can learn too by reading the bug report!)

In a few months I’ll run another of my Higher Performance Python virtual training classes, you’re most welcome to join. You’ll find details on my very-lightly-used “training email list“, you should join this if you’d like to hear about my upcoming training courses.

I make notes on some of these topics in my irregular “weekish notes” here on the blog and in my every-2-weeks “thoughts & jobs” email list. You’re welcome to join the list (your email is always kept private) if getting it in your inbox is more convenient.

At the end of my talks I always ask for a postcard “if you learned something”, I’ve just received the first for last week’s talk from the Netherlands – thanks!

Thanking Henry for sending through a lovely postcard ("enthusiastic talk, very valuable insights, best presentation" – humbled!) from the Netherlands after my talk @pydataamsterdam pic.twitter.com/jPwKU880Nz

— Ian Ozsvald (@ianozsvald) June 23, 2020

Weeknote (dtype-diet)

Data science, Week notes June 2, 2020

Over the weekend I hacked on dtype_diet – a tool for Pandas users that checks their DataFrame to see if smaller datatypes might be applicable. If so they’d offer no data loss and a reduction in RAM, for Categorical data there’s also the possibility of faster calculations. This tool makes no changes, it recommends the code you might copy into your project. I developed this as one of the ideas I didn’t get around to building whilst I was working on my book, but since that’s now published…

Talking of which – Micha and I have cleaned up our profiles on Amazon and Goodreads for High Performance Python 2nd ed. It is a pity we can’t do any book signings at events (seeing as we won’t have physical events for a while!). One thing I might do via my “thoughts & jobs” email list is to organise a Friday “coffee morning” to chat with whoever turns up about high performance Python and other subjects.

I’ve also learned that if I make a 1kg sourdough dough, immediately after kneading I can put 0.5kg into the fridge and 4 days later it cooks just fine. This is a nice refinement. I’m now interspersing buying fresh bread (for variety and convenience) with making my own and it feels like a more sustainable practice.

All going well I’ll be talking this month at PyDataAmsterdam and next month at EuroPython on higher performance Python and Pandas, I figure the dtype_diet tool will fit in nicely with a discussion about the new Pandas dtypes and their benefits over numpy dtypes.

Week(ish) note

Data science May 20, 2020

So – High Performance Python 2nd ed finally shipped (Amazon, Goodreads) – yay! In brief we’ve added notes on how you can be a “highly performant programmer”, added some more profiling, added Pandas onto NumPy, improved the Compiling to C chapter with more Numba and a new full section on GPUs (in the first edition we said – GPUs are hard, nobody uses them, avoid them – oh, how the world moved on there!), updated async, JobLib with multiprocessing, multi-core including Dask, less RAM with Pandas and NumPy and four new “lessons from the field” from very experienced colleagues (thanks to Soledad, Linda, Vincent and Valentin). I need to write some more about this.

I posted it to Twitter and got a lot of lovely feedback.

I’ve started working on a very interesting project aimed at helping the likes of the Bank of England, the UK Government, the Civil Service and more understand the impact of the pandemic on the UK economy so fixes can be designed. My involvement is just starting, I’m hoping this will grow in interesting ways.

I put a simple solution into the Kaggle Covid 19 Week 5 competition and on the private leaderboard I’m hovering around rank 10 of 90ish. My solution takes the 5 day moving average and builds wide percentile confidence bounds – it is really dumb but it turns out to be really robust too.

For now I’ve cancelled my upcoming courses, life with the pandemic continues to be weird so I’ve decided to simplify things a bit. They’ll restart in a few months.

In other news my sourdough baking continues to improve, my radishes got pretty big, my lettuce is up, the plums and pears have started to grow and generally the garden is looking awesome. I’m also working harder at training my Spaniel, we’ve got a long way to go but having her not try to flush every cat she sees in the street is a huge step forwards.

Week note

Python May 6, 2020

Well, mid-next-week note I guess. I gave another variant of my higher performance Python talk last night for PyDataUK to 250 live streamers, we had some good questions, cheers all.

On Friday Micha & I heard that the 2nd edition of our Higher Performance Python book has gone to the printers – we’d said we’d do 1 person-month each on it last summer and 9 months later (with many-person-months invested each) we’re finally there. Phew.

I now have an open PR on the dabl project to add ordinal-sorted y-axis box plot items in place of the default always-sort-by-median, which I think makes some of the exploratory process more intuitive. This also involved figuring out a new weird matplotlib rendering behaviour and writing my first unit test where I make up a dummy matplotlib figure in a test which is rendered but never displayed. There’s always a million new things to learn, right?

I’ve also been digging into Companies House data to look at how the economy is responding to the pandemic and this lets me play with some higher performance Pandas operations (for my talks) and to dig into pivot_table, pivot, groupby and crosstab along with stack & unstack in Pandas. I’ve always been confused about how many options I have available here, I’m less confused now, but I still don’t understand some of the performance differences I see for otherwise-equivalent operations. I also discovered the Pandas xs operation (take a cross-section of a dataframe) whilst reading a wikipedia page on crosstabs. Learning. Always learning.

My kneaded sourdough is improving, I’m up to 1kg now. I think I’m done with no-knead for a bit, that was fun but the really-wet dough is hard to handle. Radishes are great, pretty big now, but annoyingly the snails have found my lettuce.

Ian Ozsvald
Read my book
Oreilly High Performance Python by Micha Gorelick & Ian Ozsvald
AI Consulting

Mor Consulting Ltd. is an A.I. focused consultancy offering strategic research and development owned by Ian Ozsvald, based in London (UK).
Co-organiser

PyData London provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other.
Trending Now
1
Upcoming discussion calls for Team Structure and Buidling a Backlog for data science leads
Data science, pydata, Python
2
My first commit to Pandas
Python
3
Skinny Pandas Riding on a Rocket at PyDataGlobal 2020
Data science, pydata, Python
4
“Making Pandas Fly” at EuroPython 2020
High Performance Python Book, pydata, Python
5
Weekish notes
Data science, Python, Week notes
Tags
Aim Api Artificial Intelligence Blog Brighton Conferences Cookbook Demo Ebook Email Emily Face Detection Few Days Google High Performance Iphone Kyran Laptop Linux London Lt Map Natural Language Processing Nbsp Nltk Numpy Optical Character Recognition Pycon Python Python Mailing Python Tutorial Robots Running Santiago Seb Skiff Slides Startups Tweet Tweets Twitter Ubuntu Ups Vimeo Wikipedia

Entrepreneurial Geekiness

Weekish notes

“Making Pandas Fly” for PyDataAmsterdam 2020

Weeknote (dtype-diet)

Week(ish) note

Week note

Read my book

AI Consulting

Co-organiser

Trending Now

Navigation

Recent Posts

About Ian

Weekish notes

“Making Pandas Fly” for PyDataAmsterdam 2020

Weeknote (dtype-diet)

Week(ish) note

Week note

Read my book

AI Consulting

Co-organiser

Trending Now

Tags

Navigation

Recent Posts

About Ian