Category Archives: Uncategorized

Good management: take your hand off the saddle.

The Harvard Business Review Ideacast podcast recently had an interesting interview recently with Eric Schmidt and Jonathan Rosenberg, Google’s former SVP of Products.  The subject was lessons from Google about how to manage talent, which are documented along with other lessons in a new book the pair have written (“How Google Works”).

The main point they were making was how to successfully recruit and then make best use of technical creative people such digital product designers and developers.  The key thing that stuck out for me was the challenge of managing this kind of talent in the most effective way and it made me reflect on my own experiences of managing people.

One of the key challenges as a manager is overcoming the temptation to exert too much control and trust your team to solve the problem.  This stems from a natural concern for quality and timeliness, which are clearly desirable concerns to have. However, the temptation is then to over-specify the outcome you want in order to delivery high quality on time.  What this then often leads to is your team either constantly having to check back in with you to find out if they are going in the right direction, or delivering something which doesn’t fit in with your specific vision. They then get disenchanted and demotivated and the project gets delayed or derailed.

A much more effective way to delegate to your team is give them the problem, rather than dictating a solution you have already come up with.  In giving them the problem, you also have to give them all the knowledge you have that is relevant to it – who are the key people to talk to, what are their motivations, what data will be useful, what are the key risks with it.  How much knowledge you give them is obviously determined by their experience, how important/risky the project is, and how much you want them to have to find out about things for themselves.  However, the simple principle of giving them the problem, rather than the solution will have many benefits with the main one being that the outcome will be the result of input from the whole team rather than just the manager.  If you have have put a lot of effort into recruiting talent, then this is clearly what you want!

My most rewarding example of this was when I was able to hand over the writing of a ministerial submission to a student who had been working for me for the last year.  By giving her as much information as I knew about the situation and the freedom to solve the problem in her own way she was able to effectively communicate the issue to our minister with very little input from me.

Size doesn’t matter

…at least when it comes to information!

As the statistical nerds amongst you may be aware, there is a bit of a backlash going on against “big data”, fueled partly by discovery of the hubris in Google’s attempts to predict flu outbreaks. Tim Harford is thankfully on the front line, as well as the Economist.   Health sector colleagues have also focused on the particular limitations of the application of “big data” philosophy to health.

The conclusions that these critics have come to is one that those who have any knowledge of statistics had probably already drawn: apart from for some very specific purposes, there is very little to be gained using “big” datasets. Once you have a good conceptual understanding of what is generating your data, the value of an additional data point drops exponentially after the first 10 or 20.

This can be demonstrated using an example from physics. Start off by talking to your friendly neighbourhood physicist, and she will tell you that there is a law to explain the relationship between the temperature of a strip of copper and its length. Armed with this knowledge you can then perform an experiment to test it, observing a strip of metal at different temperatures. This will give you a graph that might look like this:

temp1

With only 10 data points it is trivial to verify the prediction of how copper will expand. Another 10 is not going to change your conclusion, nor make you much more certain about it:

temp2

Much of the noise about big data has come from the IT industry, for whom big data does present some non-trivial problems. For example, there is the oft-mentioned fact that Rolls-Royce jet engines generate hundress of gbs of data every second. Getting this data off a plane, stored somewhere and fed through some system for detecting malfunctions is a real feat of computing.

I’m not going to reiterate the flaws in unconditional claims of boosters of Big Data as others mentioned above have done this much more eloquently. My plea is much more practical and relates to something we all have a personal interest in.

The scandal surrounding care.data  quite reasonably frightened a lot of people. Most people in the UK view their medical records as very personal information and the way the potential re-use of this information was presented left a lot to be desired.

However, the real tragedy of this episode was that it delayed for years an innovation that could be the most powerful force for improving the effectiveness of the NHS, and reducing its costs. This innovation is linked data. It is not a complicated idea, just that of joining up information so that information about how someone is treated in one place is joined up to how they are treated in another.

As described in the article in the Health Services Journal by Axel Heitmueller and Sandy Pentland (linked above, and again), joining up multiple data sets, in particular across different care settings (for example acute hospitals and community providers) brings many more benefits than ‘big’ data as commonly shouted about.  The aim of Care.data was really to facilitate medical research.  However, for the NHS at the moment it is much more important that current treatments are delivered more efficiently.  Rather than glamorous product innovation to create new treatments, this means process innovation: making things work better.

The primary benefit of linked data is in helping healthcare providers and regulators better understand patterns of healthcare and provide a more seamless journey for us, the patients.  Everyone has an incentive for this to happen – patients obviously have a better time if the people caring for them can talk to each other and provide a smooth journey between services, and as documented by the vast literature on Lean, providers invariably save money when customers/patients have a better journey.

Rather than being distracted on the one hand by silly market fads about big data, and on the other by the merits of sharing medical data, let’s all demand something that undeniably benefits patients and also has the potential to save the NHS a lot of money.

Here is a message to go and repeat wherever you can – “link my data!”

The Curve, by Nicholas Lovell

Wandering around the Whitechapel library last week I was attracted to the vibrant red cover of this book, along with the fact it had a blurb by David Rowan, the editor of Wired UK.  The slightly less vibrant content of the book is a hubristically contrived framework for thinking about how to successfully market goods and services in a world where so much is increasingly available for free.  The “curve” in question is the demand curve, that helpful tool from first year undergrad economics courses which illustrates how the quantity of a good that is demanded changes as the price changes.  For most people, if the price of something falls they are more likely to buy it.

Mr. Lovell’s point is that most firms face a demand curve where there is a small volume of people who are willing to pay a lot of money for their stuff, and a large volume who are not willing to pay very much.  The globalisation of competition means that in many cases the price of things has fallen to the cost of production.  Firms therefore have to use new methods in order to either a) identify those people who are willing to pay more b) move people up the curve or c) use the sale of zero-profit goods to drive sales of higher margin goods.

He comments on various business models that have either emerged in response to technological disruption, or have seen new or more widespread use, for example, freemium in the case of music distribution and different use of platforms/two-sided markets by Apple with its App Store and Amazon with its Kindle.

However, it’s really just a number of extended magazine articles which don’t hang together.  There isn’t a cohesive story, other than that technology is disruptive.  I don’t need to read any book to teach me that lesson, hence why I’ve only skimmed through about 20% of this one! (For a lesson in how to write a compelling book on changes in global trade and marketing, read Thomas Friedman’s The World Is Flat.)

If you’ve been living in a cave for the past 15 years and haven’t read anything in the business, technology or popular press about how “times they are a changing”, then this is definitely worth your time. Otherwise, give it a miss.

Easy answer…to an easy question

I previous posts I have mentioned that I was going to try and look at forecasting the volume of treatments of admitted patients carried out in English hospitals.  The graphs below show first the national volume of treatments, across all hospitals. (FCE stands for finished consultant episode, the unit treatments are counted in).  The second graph shows a zoom in of the results of the forecasting methods.  If you are wondering, yes methods 3 and 5 have produced almost exactly the same result.

Image

 

Image

It would seem from this graph that method 2 is the most accurate. Method 1 was never going to be any good: it was really just me checking I was using Stata correctly to create the new forecast data.  Methods x-y are all based on econometric estimators.  Although they appear to all present the same results at a national level, at the level of individual hospitals their accuracy does vary slightly.

The statistics used for more accurately measuring the forecasts at the level of each hospital (which is what I care about) are based on the average difference between the forecast value and the actual value.  Expressed as a percentage the econometric methods had an average error of between 0.5% and 0.7%.  This is not not too shabby.

Method 2, which says the growth rate will be the same as in the last year observed proves to be the most accurate at the level of individual hospitals, with a mean error of 0.01%: the most successful by a long way!

The reason why this task an “easy question” is because while the data are complicated – many observations across multiple hospitals – forecasting two data points when the data do not vary a great deal from year to year means that any reasonable method is never going to be significantly wrong.  What I might do next is look at some hospitals in more detail, possibly those with the worst forecast, and see if there is anything they have in common.

It would also be interesting to try forecasting over a longer period of time.  Another option is to download the monthly version of this activity data and use that, which would mean twelve times more detail! I could also use Monte Carlo simulation (doing the forecast lots of times) to get a distribution of results, rather than just a single point estimate for each year.

Dipping a toe in the water

Before reading this, it is worth skimming my previous post on the topic of forecasting English hospital admissions.

Mean growth per year

Mean growth of each trust

The two graphs above show the results of my initial pokings into the NHS hospital activity data.  The first one shows quite clearly how activity has in average increased every year. (I’m not yet quite sure why it isn’t showing any data for years 0 and 1.  By rights it shouldn’t show 0, but 1…)

The second shows that while almost all trusts have positive mean growth over the period, there is a large minority which have only experienced very modest positive growth, much less than the national means in the first graph would imply.

While mildly interesting on its own, this has implications for the inferential analysis which is going to follow.  One of the key issues in performing econometric analysis with panel data is how you treat your units, in this case hospital trusts.  Under one approach, you assume that that each unit has its own unique effect on the variable you are analysing, but that these effects are random.  The second approach says that they are not random but driven by some systematic differences in the units.

Based on intuition one would have thought that the random approach would not be appropriate for hospital trusts because the growth in activity is going to largely be driven by their local population and the the funding levels of their local Strategic Health Authority, i.e. there are systematic differences.  The second graph doesn’t really help us decide which approach is more appropriate because it shows the trusts as being quite neatly distributed, even if the mean is skewed by some outliers.

This means that we will have to use statistical tests to decide which approach is better, and possibly just see which makes the better forecast.

NHS Hospital Activity – Looking Under The Hood

So today I started a little project to look in to how the amount activity performed by NHS providers changes over time.  The focus is going to be on the aggregate amount of activity performed by each individual provider.  There are potentially a number of stories in these data but the most interesting is whether it is possible to forecast the amount of activity for the next year.

Today I have started building my dataset from the annual activity datasets published by the Health and Social Care Information Centre (HSCIC).  I was hoping that the collection of the data into a pleasing balanced panel would be smooth going.  How wrong I was.

Firstly, one would have thought that in this age of Data.gov.uk  one of the largest holders of publicly accessible data in the UK would have a sophisticated system for storing and searching through their data.  The huge value of the data that has been fully uploaded to Data.gov.uk is that it is very easy to find different years etc. of the same data, or different categories.  Like the Office for National Statistics, the HSCIC has chosen to take its time in systematising the storage of its data.  It is possible to search for the Hospital Episode Statistics on Data.gov.uk, but you will not find any spreadsheets, let alone tidy csv files.  All you will find is links to HSCIC pages.  Welcome to the 90s!

Once you’ve resigned yourself to trawling through the HSCIC pages you will encounter multiple frustrations with corralling together time-series data: each year’s publication has it’s own page with inconsistent titling so they don’t all come up together in a search; the spreadsheet for each year’s data has a different structure so you can’t pull it out with a script and the same variables have different names in different years.  It’s almost like someone is intentionally trying to make it hard to do anything with this data…

For anyone reading this with knowledge of these things, I have only been working on data for Admitted Patient Care (APC).  The other big category of hospital activity is Outpatient Procedures (OpProc).  The major difference for these purposes is that the activity measured is different.  For APC, the metric is Finished Consultant Episodes,   whereas in OpProc measures procedures.  For any sort of inferential model of hospital activity it will be necessary to look at both because for some conditions a patient can be admitted or treated as an outpatient.

More to follow as I stick my head further under the hood!