Tuesday, July 14, 2020

California's Goober Doubles Down... I Do, Too.


Background and Bottom Line


Our pitiful excuse for a Goober doubled down on his recent 'rollbacks' due to increases in CoViD-19 cases (actually on his twitter feed he wrote "spread at alarming rates"). I made my case last week (here) for why I thought it was a mistake when he rolled back San Diego last week, and now he's doubling down. Well, two can play at that game.

I have a lot to say, but for those of you who don't want to read the whole thing, let me give the most important data right up front. I predicted the death totals for three consecutive weeks. Let's call them weeks 0, 1, and 2, since those were the numbers of weeks into the future the prediction was for. OK, predicting "0" weeks into the future isn't really a prediction, but it does help to validate the basis for the other predictions. So how have I done?

1) For week 0 (week ending 7/4/2020): I predicted, 33 deaths, and there were 27 observed deaths. Not bad. The prediction was 20% on the high side. This could be due to using data over the entire historical time to predict what is happening "Now". If the treatments were improving, or if there are simply far more tests finding a greater fraction of the cases (better/more complete contact tracing), then you have to expect an over estimation.

 2) For week 1 (week ending 7/11/2020) I predicted 66 deaths. If we believe the 20% over estimation, maybe it should be 53. In any case the actual number of observed deaths for that week was 35. So actually it begins to look like I may be significantly OVER estimating the deaths. So when I claimed the Goober was imposing reopening rollbacks based on faulty metrics, I appear (at least so far) to be correct. Now he's increasing the rollbacks! Yet the raw prediction for the total for week 2 is still 61 (perhaps 49, maybe less), and we will see how that comes out.

Without bothering to 'show my work', I will state that I can now make a prediction for week 3 (week ending 7/25). It is 62.13, which I will round to 62. So there is still no prediction for a major rise in deaths and the Goober is reacting to a completely misleading metric (at least for San Diego County). 

So far the score is: Goober: 0, Me: 1.

A More Thorough Analysis of the Prediction Methodology and Assumptions 


If you go through my long post from last week you will see that I based my predictions on rates determined by dividing the deaths by the number of cases two weeks previously. That was of course a completely ad hoc assumption based on the reports that many deaths take 'weeks'. During this past week, I began to wonder, is that the right delay? Can I actually find an answer in the data?

KT was kind enough in the comments to the post last week to provide a pointer to Gummi Bear's Twitter thread where they talked about some observations they had made of CoViD data. In this there were lots of numbers and more than a few plots. The plot that struck me was a side by side comparison of deaths over time for New York City and Spain. Even though I had seen that plot for NYC before, the fact that the plots looked nearly identical triggered something.

Some time ago I had noticed (on the very excellent NYC CoViD site) that the curves for cases, hospitalizations, and deaths all had the same shape, but with the deaths clearly lagging the other two by about a week (don't just believe me: go to the site and scroll down to the "Daily Counts", it comes up on the Cases data, click on the Hospitalizations tab, see that it changes only modestly, then click on the Deaths tab, except for shifting to the right, it also doesn't really change). Anyways, that flash of memory suggested to me that the proper lag might be only 1 week, despite the stories of folks lingering.

Just trying to plot the raw daily data, with all the inherent noise seemed a fool's errand, but I decided that if I averaged the cases over 7 days, and the deaths over the same 7 days, I ought to be able to then plot the deaths vs cases, where if I delayed the cases by an optimal number of days I should be able to make a reasonable straight line. So off I went. To 'cut to the case', the optimal number, was, to my great surprise, four (4) days (see the plot below).


Plot of  7-day averages for deaths vs cases (delayed by 4 days) for New York City.  The quality of the correlation is obvious.

Now, I concede that the data, particularly back in time, may well be less than directly comparable to today (that is, of course, exactly what I claim our Goober is doing, so I must be careful not to fall into the same trap). How might it be easily different and in a way that would effect the apparent lag time? One very obvious way (well obvious to some - Thanks, KT!) is that back then in NYC, the place was suffering from massive numbers of cases, and insufficient testing availability. If that meant that people were not getting tested as soon as they would today (likely!), then it would reduce the apparent lag versus what we'd observe today. So really I need to use the latest and best data I can get my hands on. For me that's the San Diego County data. While I've been saving lots of case data for the county for almost 4 months, I did not start to collect the daily demographics for deaths until this month. So I won't have sufficient data to say anything more definite for another week or two. I promise to do a follow on once I do.

Fool Me Once, Shame On You. Fool Me Twice, Shame On Me 


In the previous post, and so far in this one, I have focused on only one metric. But let's take a quick look at all the metrics that San Diego County monitors, where we have gone into "alert" mode. The first is the same idiotic metric that our Goober is apoplectic about, namely that the number of cases is more than 100/100,000. The other two county metrics are "Case Investigations" and "Community Outbreaks".

The Case Investigations metric is indicating that the county isn't getting the contact tracing started fast enough. That's purely a function of the number of people doing the tracing and the number of cases. Clearly this is a reasonable thing to watch, and it is important to trace as fast as possible (to try to get out in front of the spread), but it isn't really contributing to our reclosing. The other metric, Community Outbreaks, is much like the case numbers in that the county is not comparing apples to apples.  Whether the metric as formed (more than 7 outbreaks in 7 days) is reasonable, depends. For example they are defining an outbreak as 3 or more people, who are from different households, that test positive, and are associated to a specific location/event. A week ago the county reported 21 such outbreaks, 15 of which were associated to "restaurants/bars". Yesterday they reported 17 outbreaks, with 7 associated with restaurants/bars. So far, so good. A reasonable person should ask two questions: 1) Are they truly investigating ALL the possible sources of interactions? and 2) What would we expect from random chance of people who were already infected being in the same place at 'the same time'? I really want to tackle the second in some detail, but that distracts from the far more significant first question. Consequently, I have moved the entire discussion of the second question to the bottom of this post under the heading "Real Effect or Random Correlation?".

But we still have the first question I posed above to deal with: Are they truly investigating ALL the possible sources of interactions? Here I can give an absolute answer. NO. Actually, Hell No! Here the county (and state) have completely lost their moral compass and intentionally failed to perform their duty. They should all be be charged with malfeasance and fired.

But what could cause me to be so sure and so angry? Simple. They refused to even ask if any of the people testing positive attended any of the protests (a.k.a. the riots). As we know these were also attended by predominately the same age groups as were also frequenting restaurants and bars. However, due to the intentional malfeasance of the county health officials, we can never know for certain if these played a role in the spread.  We can however demonstrate a smoking gun...

I can say definitively that the jump in cases is observed as having occurred during the 10 day period June 14 - 24. You can convince yourself of the same thing by merely looking at the plot below of positive test cases vs date reported.
Plot of the 3 day running average of the number of new cases versus date for San Diego Couny.  The two red points represent the data for July 15 and July 25, in between which the entire rise of case numbers from about 120 to about 470 occurs.

The protests occurred mostly during the period June 1 - 13 (I did a web search for "protests San Diego" and just noting the dates of the reports). If we assume these were significant vectors for spread, we would have seen a rise that was most pronounced in precisely the date range we see it (in the plot above, the two red points are for 6/15/2020 and 6/25/2020). If the rise were due to restaurants/bars the rise would have a coincident rise point, but it would not have leveled off so abruptly, while an abrupt level off is completely reasonable for the end of the protests. In fact, had the cause truly been the restaurants and bars, the earliest we would have expected it to level off, assuming the Goober's reclosing of the bars and restaurants was the primary factor, would be about a week after it took effect, which would work out to... today, the 14th of July.

I see no way to argue anything other than: 1) the 'protests' were the primary driver of the case increase in San Diego County, 2) our Goober and County officials are punishing business that had, at most, a minor effect on the case rise, 3) to increase the closures, in San Diego County, due to disease "spread at alarming rates", is unjustifiable, alarmist, and borderline criminal, and 4) every government official who encouraged, or even merely acquiesced to, not asking about attendance at the 'protests' should be fired immediately for the damage to the people and business of San Diego County they caused by not properly doing their jobs.

Real Effect or Random Correlation?


I have tried to gather the best possible data to answer this question. So here goes, and please bear with me. After doing a bunch of web searches I have concluded that there were well over 180 restaurants before the first lockdown. That is the number of members in the San Diego County Restaurant Association (listed as "180+"). I do not know if a restaurant group (all owned by one person/entity) count as 1 or many. Clearly the number might be bigger, possibly way bigger. I'm also ignoring the number of non-restaurant bars (and that's a bunch too, but I can't really get a good handle on the number). I also don't know how close to capacity the restaurants have been during the reopened period (the news stories I've seen suggest that they've been near their capacity). I also can't say what the true capacity would be. I also don't know how close in time the people needed to be to count as 'at the same time' in the contact tracing. But let's take some educated guesses.

I know, from personal observation, that most of the restaurants I used to frequent would have signs listing their capacity. Those listed capacities were always over 100, usually more like 120. During the reopen period there were requirements to maintain social distancing, so the density was clearly reduced. I also know from the local media that many restaurants had been using outside spaces to add some of the capacity back. So let's take a swag at something like 50 being a reasonable estimate of the average capacity (ASIDE: This is a really important number as the more people there are, the higher the likelihood of random clusters, i.e., groups of 3).

We know that restaurants generally don't seat every table, then later clear out all tables and then reseat every seat and repeat. It is much more continuous than that. Let's guess that during the lunchtime, something like 100 people can be there close enough in time to be counted as 'at the same time'. I'd guess that dinnertime you'd see something larger, maybe two groups of 100. We also know, from the plot above, that the county has had a variable, but consistent case count over the last two weeks of between 400 and 500 (the daily average over the last two weeks is actually 471). We could just use this, however, when we start talking about the number that might be going out to restaurants it is likely to be too big. That '471' is all age demographics summed up. But we know that it is almost entirely the under 40 crowd, and truthfully almost entirely the 'over 20 and under 40' crowd that are going out to the restaurants. So let's just count the cases for those aged 20 - 40. The numbers are that over the last two weeks there were 3,296 cases in this age bracket, and the county population is (for the same age group) about 1,046,000. If we use these data to determine the number of active cases on any given day (by multiplying by an average duration of 10 days and then dividing by 2 to remove the identified cases - who presumably stop going out) we get 1,177 cases (infected but not identified) in a population of 1,046,000, or an average infection rate of essentially 0.11%.

Now we can calculate the expectation of finding at least 3 random people in a 'seating' of 100 as 0.0013. If we then take 3 'seatings' per day per restaurant and 180 restaurants we expect to see about 0.61 'outbreaks'/day (or 4 per week). Note these are 'outbreaks' where the association is not that they spread the disease while at the restaurant/bar, but rather that they just happened to randomly be in the 'same place at the same time'. I strongly believe that this estimate is a severe underestimation. I believe this because I expect that the cases will be concentrated in the fraction of the population that are going out (for example, I have 3 sons in this age group and NONE of them, their wives, and certainly not their children, are going out). I would think that I could easily be a factor of 2 or more low on the expected number of random 'clusters'. If so, then ALL the observed restaurant/bars 'outbreaks' of last week are actually just non-causal clusters. If the factor of 2 is correct, then there may be essentially no true Community Outbreaks, and even if 3 are real 'outbreaks', then it gets closer and closer to having the actual number of outbreaks to be below the metric. Anyways, yet again misuse of a metric is so easy to see, if you bother to look...

Tuesday, July 7, 2020

Figures Don't Lie, But Liars Can Figure


Updated on 7/8/2020 - see below

And sadly I must report that the San Diego County Health Officials (and the CA State folks, up to and including Goober Newsom) have been 'figuring' (or perhaps just showing their ignorance).

First a bit of background.  When Goober Newsom allowed the various counties around CA to 'open up' (albeit slowly), he put some 'metrics' in place to monitor the counties behavior/results.  Now at first blush this seems a reasonable and responsible action.  However, the metrics (set in mid April) have not been modified to account for the changing reality, and as a result are now being applied, I believe improperly, to shut businesses down in San Diego County (and I must wonder if in other counties as well).  I have been saving lots of the historical CoViD data for San Diego County and I dug into it to determine if they are comparing apples to oranges.  And the result is (drum roll)...

They are comparing apples to squirrels.  The data today simply can not be compared in a straight forward way to metrics based on data from back then for a myriad of reasons.  I will now attempt to convince you, dear reader, that I know what I'm talking about.

The metric which has caused us to go on the Goober's 'watchlist' is that we have exceeded the allowable average daily count of new cases (a 14 day running average which is also divided by the county's population - in 100,000's, for San Diego that means it is divided by about 33).  Now if that is unclear, please excuse me, but that's the metric.  When the metric exceeds 100 for 3 consecutive days, the state starts to 'watch' and if it stays above 100 for 3 more days, then the state orders all 'indoor' business to close for at least 3 weeks.  But why?  If we are aiming to someday reach herd immunity (and so long as we lack a vaccine, that's the only path available to a general return to normalcy) then having the cases rise speeds the process.  Of course I would prefer we do this as safely as is reasonably possible, so it is reasonable and prudent to try to keep cases, hospitalizations, and ICU admissions down, but really we should be striving to keep deaths down.  You will note that the metric simply does not take this into account.  Additionally, the metric does not take into account the tremendous rise in testing since those early days.  Perhaps they were simply trying to get a surrogate for deaths, not after people had died, but in metric that would predict deaths.  At first blush cases seem a reasonable choice.

As of the day they set the metric, the death rate was running at about 6% in San Diego County*.  At the time the highest number of deaths in a week was 40.  If we divide that by 33 to get it per 100,000, we get 1.2 deaths per 100,000 per week, at the maximum.  But the cases were running at about 570 per week, and that's about 17 per week per 100,000.  So let's work with the assumption that they were prepared to accept the deaths that would be produced if the case rate rose to the 'magic' 100 per week per 100,000.  That works out to about 1.2*100/17 or about 7 per day per 100,000.  So far, while the metric is convoluted, and only marginally predictive, it would be not insensible.

But how does that apply to today?  Somewhere between poorly and utterly and completely irrelevant.

Why?  Because the demographics of those getting the disease has changed markedly, the testing capacity has seen a major growth, and, I suspect, the treatments are better.  Let's agree to ignore the last possibility and focus on the first two.  As to the number of tests, back in mid-April we were running around 1,100 tests per day in San Diego County.  Today that average is more like 6,700.  If the actual fraction of people infected stays the same, and the number of cases is substantially larger than the number confirmed, then the total testing positive would be a fairly constant ratio of the tests, so we'd expect to see something like 6 times the number of cases (and it is  a factor of 4.5, not too far off). Additionally, as contact tracing has improved we'd expect to see even more of the asymptomatic cases be identified, so just using case numbers is nonsense! And we haven't even started on the issue of the demographic change.

The demographics of the cases for San Diego County in mid-April consistently showed that the three most susceptible age groups, 60-70, 70-80, and 80+ accounted for 25% of the cases, and nearly 90% of the deaths (with the under 40 demographics accounting for about 35% of the cases, but only 2% of the deaths).  So if we really want to keep the deaths down, we need to worry about the 60 and over crowd (and be somewhat concerned as to those 40-60).

Well we have the data, so what's the death rate today and what do we expect to see for the next 2 weeks (based on today's cases). As of the last three days we have seen a total of 0 deaths.  OK, even I don't believe that.  It's probably something of an artifact, likely caused by the holiday weekend.  Let's look instead at the deaths for the 7 days prior.  Total deaths: 27.  After we divide by 33 (to get in per 100,000) the total is less than 1 (it is just under 0.79).  Not even as high as the totals we'd seen before, let alone jumping up towards 7.  How about the projection for two weeks out (based on the fact that there were 3392 cases reported in the week 6/28-07/04)?  Well we need to break these down into their demographic groups.  Fortunately, I have that data (it is in Table 1 below for those who want to check my math).  There it is, the projected deaths per week per 100,000 in two weeks is likely to be less than 2.  Again, no where near the 7 they seemed to be willing to tolerate.  Given that, the closure of businesses is completely unwarranted.

             
UPDATE: Note added as proof.

It strikes me that if my method of predicting deaths from cases has any validity I should be able to take the case data from two weeks ago and predict the deaths as of today.  I can't believe I didn't think to do this yesterday.  Anyways the data are shown in Table 2 below.  The prediction is a total of 32.74 deaths for the week ending July 4.  The observation was 27.  Anyways that seems close enough to me.  It also suggests that prediction for 2 weeks from now is likely to be high by about 20% (this could be due to lots of things, but I'd bet on most of it being due to more cases that are asymptomatic being found through contact tracing, which would have simply gone as undetected previously, increasing the case count, but not contributing to the death total).  I won't be surprised to see deaths for two weeks from now to come in at around 50 (or about 1.5 per 100,000).  I also added Table 3, which is the prediction for the deaths for the week ending 7/11/2020.  (Again, I can't see why I didn't think to do this earlier.)  So we'll see how this works. Oh, and that's for data ending before San Diego got on our Goober's Watchlist and it is slightly higher than the prediction for deaths for the week that cased us to go on the list, but still essentially 2 / 100,000.

             

* - this number is what I get by taking an average number of new cases a little before the date the metric was set and an average number of deaths over the time period 14 days after that.  It's crude, but it isn't unreasonable.

             

TABLE 1: The death rate for each age group is calculated by taking the number of deaths (as of 7/4/2020) dividing those by the number of cases two weeks previous to that (as of 6/20/2020).  The new case totals are for the week ending 7/4/2020, so the death prediction is for the week ending 7/18/2020. The total number of deaths would be expected at about 61 for that week (compared to about 40 above).  That's still less than 2 / 100,000.

Age group New cases  Death rate   Expected Deaths 
0-10
110
0/254
0
10-20
267
0/558
0
20-30
1109
3/2152
1.55
30-40
674
4/2000
1.35
40-50
446
12/1633
3.28
50-60
381
29/1810
6.10
60-70
235
60/1129
12.49
70-80
93
93/645
13.41
80+
74
186/604
22.79
 Total Expected deaths 
 60.97 ≅ 61 

             

TABLE 2: The death rate for each age group is as in Table 1.  The total cases by age group is as of the 7 days ending 6/20/2020.  The expected deaths are for the week ending 7/4/2020.  The total number of deaths would have been expected at about 33 for the week (compared to 27 observed).

Age groupNew cases Death rate  Expected Deaths 
0-10
48
0/254
0
10-20
122
0/558
0
20-30
360
3/2152
0.50
30-40
221
4/2000
0.44
40-50
173
12/1633
1.27
50-60
207
29/1810
3.32
60-70
125
60/1129
6.64
70-80
53
93/645
7.64
80+
125
186/604
12.93
 Total Expected deaths 
 32.74 ≅ 33 
 Total Observed deaths 
 27 

             

TABLE 3: The death rate for each age group is as in Table 1.  The total cases by age group is as of the 7 days ending 6/27/2020.  The expected deaths are for the week ending 7/11/2020. 

Age groupNew cases Death rate  Expected Deaths 
0-10
83
0/254
0
10-20
210
0/558
0
20-30
734
3/2152
1.02
30-40
483
4/2000
0.97
40-50
341
12/1633
2.51
50-60
269
29/1810
4.31
60-70
201
60/1129
10.68
70-80
116
93/645
16.73
80+
98
186/604
30.18
 Total Expected deaths 
 66.40 ≅ 66 

Friday, May 8, 2020

An Update and Some Thoughts About the CoViD-19 Situation


So it's been about 2 weeks since I last updated about where I see the whole CoViD-19 situation going.  This is due to my wanting to have more data, before I started ranting.  Well much of that data is in, so here we go...

First, we've seen some of the most amazing things with respect to reopening.  Things I would not have thought would happen in my lifetime. First, we've seen people openly carrying weapons on to the grounds of the state capital building in Lansing, MI.  (Interestingly, they even carried them into the building lobby, but were not allowed to carry in their signs.  That's just plain surreal.)  We've seen people starting to openly defy the stay closed orders, some for beaches in Orange County, CA, and now more and more examples of business owners who are reopening their business without 'permission'.  The people are speaking and voting with their feet (and "arms").  Plus many "red" Governors are rapidly opening their states. The map of open versus closed states is looking more and more like the electoral college results of the last election.  So who is right?  The people?  The "red" Governors?  The "blue" Governors?  Some of all?  None of the above?

Sadly, I am forced to conclude that the correct answer is 'none of the above'.  Here's a smattering of reasons why (focusing on states I have a personal interest in):

1) Not one group is using the only sane and sensible method to reopen.  The one I blogged about back on April 19 (that's 19 days ago!).  The one based on the data.  Nope, not one of all the Governors, regardless of their party affiliation, are actually using the data to do things right.  I have yet to hear a single Governor talk about opening the economy, but warning the old or those with underlying conditions to remain as isolated as reasonably possible, and to tell the rest that they simply must stay away from those who are still isolating.

2) Ohio: Even DeWine, (a 'red'), who has done, so far I can tell, a wonderful job to this point, is simply ignoring the data and using the broad brush approach.  It's like to trying to fight a fire by spraying water in random directions, even if the blaze is not within view.  As an aside he's fallen into the trap I pointed out back on April 26.  He's quoted as saying:

[The number from Monday, April 27] of 362 more cases was below the recent five-day average of 442 cases. He pointed out hospitalizations of 54 on Monday as compared to the daily average of 70. “We’re not going down [my emphasis] ... but we’re moving in the right direction,” he said. Before providing the reopening details, DeWine again stressed that an increase in virus testing and contact tracing to track down those potentially exposed to coronavirus will accompany the restarting of Ohio. (see here)
Clearly he 1) can't see that increased testing will push the reported numbers up, even as the rate of infection drops, and 2) seems unable to recognize that 362 is actually less than than 442 and 54 is less than 70.

3.A) California, state: Our wanna be socialist tyrant (yes, a 'blue') has had his plans thoroughly changed due to pressure from the populace.  First, he wanted to re-close all the beaches, but had to back down in response to the overwhelming backlash.  Then he shifted from the first phase of reopening going from "weeks away" to "days away" when the people started realizing that he had no intention of ever giving up his total control and just started reopening on their own.  But he's also still pushing stupidity over data.   He isn't making the obvious warnings, but he is requiring that to reopen, every business has to take the temperature of every employee as they report for work.  Does he really believe that's anything other than a freaking waste of time?  The number of people who have the disease while remaining asymptomatic is not clear, but it is certainly much greater than 0.  I've seen three reasonable estimates.  The pregnant women checking in to a New York hospital for labor that tested positive were 33 in 215.  But only 4 had "symptoms", that's 88% asymptomatic.  The USS Theodore Roosevelt reported that "about 60% of the people who tested positive had no symptoms."  Finally, the reports from a Tennessee prison are that "98% of the inmates and staff who tested positive did not have any symptoms".  Take the middle value.  Only 1 in 9 will have a fever.  So go ahead and require temperatures be taken.  It won't hurt, but it isn't even in the ballpark of being a reasonable way to identify the infected and really it strikes me as possibly illegal (I'm not a lawyer, but I can read, and I read the HIPPA law).

3.B) California, San Diego City/County (a mix of red and blue - how refreshing!):  They have a lot of good data, on a very accessible website.  They've just starting adding a new plot to their data.  Essentially one of the ones I showed in my post of the 26th.  It's a plot of the running average of the fraction of tests that come back positive (here).  It is wonderful.  It shows that the rate of positive results is slowly declining (and has been since the beginning of the plot back on March 26th).  Initially, it was about 7.5%, but has rather steadily declined to about 6% now.  That is as close to proof that the county is showing progress in controlling the rate of spread, as anything could be.  But can we reopen?  Yes, but only very slowly, because our Governor says so.   Oh, and nothing about it might be smart to have the young healthy open faster than the old.

4) Maine:  The Governor (blue), is saying “We are thinking about making some changes before Memorial Day,” (Yes, that's basically a month from now).  So how horrible are the conditions in ME?  Well of their population of 1.33+ million they have 37 people in the hospital.  Yes, that's the stunningly high rate of 0.0028%.  Compare: New York City had, at the peak, about a month ago, around 1,600 people added each day to the hospitals.  So with roughly 6.5 times the population they had 43 times as many hospitalization each day than ME's current total number.  If the average hospital stay is a week the number in the hospitals would have been just over 300 times as big. The reported rate of current active cases in ME is pretty stable at about 430.  New York City, at the peak, was reporting well in excess of 5,000 each day (if the infection lasts 10 days that's a stable rate above 50,000).  Given the deep and overwhelming concern Gov Mills is showing, and given that they have about 140 traffic deaths each year, I can't see how she will sleep at night until she is able to ban cars.

5) Michigan (blue): The story here is so crazy I don't even know where to begin.  Let's just agree that any set of policies that results in significant numbers of people walking on government property carrying semi-automatic weapons is probably worth re-thinking.

6) New York, State and City (both blue as blue can be):  They have a nice website for the City.  They have some really good data (even if it keeps changing and is really only trustworthy once the data gets to be about 6 to 10 days old).  The bottom line is that the number of cases is dropping rapidly.  Doing a quick fit of the data (Mon-Fri average new cases) to an exponential decay indicates that for the last three weeks their cases have been dropping by a factor of ~3/4 each week, presumably while the number of tests has been growing (that's an assumption on my part, but I can't believe the number of tests is dropping).  The numbers of hospitalizations shows the same shape.  So does the number of deaths (although that lags by about a week).  To really believe that those are real (and I do) you have to conclude that a large fraction of the entire city has had the virus and what we are seeing now is the onset of herd immunity.  [Aside: The city reports 176,000 confirmed cases.  At the rate they are going by the end of May it will be around 200,000.  We know that the actual number of cases is much larger than the confirmed cases.  The measured/estimated ratio runs from 16 to 190.  The one that can be calculated from the one NYC 'experiment' is 32.  That suggests that 6.4 million of the total population of 8.6 million has been infected.  That's about a 75% infection rate for the population.  That would agree that herd immunity is just around the corner for the city.] We've seen them send the USNS Comfort back.  Clearly things are much better than a month ago.  So is the city preparing to reopen?  Not anytime soon.  I don't get it.  If the data suggest that anywhere can reopen it is New York City.  Sure they should keep the social distancing and increased hygiene going (it really would be just stupid to not do that).  Yes, of course, only the young and healthy should open, everyone else needs to remain isolated.  But seriously, anyone who looks at the numbers knows that NYC can reopen, and reopen now.

How about the state?  Well Gov. Coumo is considering starting the reopening in upstate New York around May 15, if they can meet his 7 criteria (here).  But the rates of infection there are not what they are in NY City.  The need for smart reopening in the upstate sections are pretty much as for any other state.  But not a peep about about the need to be, and how to be, smart. Well not exactly.  Not a word about it from the Governor, or anyone at the City, but there is this Issue Brief.  It seems familiar to me somehow.  But no one is listening.

So that's my reasoning for concluding that not a single state will reopen smartly.  So if we are doomed to open stupidly, what should we expect?  We should expect the number of cases to increase.  We should expect the number of deaths to increase.  We should expect that the politicians will all blame each other for not following and not enforcing the (patently stupid) guidelines.  We should expect to be forced to endure "Shutdown: The Sequel".   We should expect that after several more months of forced confinement, we will go at round two of the reopening.  That will go a bit better.  But then the fall and winter will set in, and the virus will come back again.  There still won't be a vaccine.  We still won't have developed anything even approximating herd immunity (I'd guess that the actual immunity will be about 20-25%).   Oh, and the economy will be totally destroyed.

We still have a chance to develop the herd immunity.  It will require that the young and healthy go out and get sick (but mostly just a little sick, many won't even know it).

I certainly hope that wasn't as depressing for you to read, as it was for me to write.

Sunday, April 26, 2020

Will They Ever Let Us Reopen the Country?

Let start with the answer: I have grave doubts that any of the blue governors will reopen their states in the next two months.  I truly hope that this is not due to them wanting the economy to completely collapse (and thereby hurting the president's chances for re-election), but I do not believe they are above such heartless calculations (and who cares how many people that might kill, they would still blame the president).  But now, with the political machinations covered, let's talk about how they can (will?) rationalize such policies.

First,  there are two conditions 'needed to reopen', a 14 day period of a downward trend in the confirmed cases and sufficient testing to be able to track and test.  I believe these two conditions are diametrically opposed.  The reason I believe this is that with increasing testing comes increasing positive tests (you see what you measure).  Can we demonstrate this?   Well let's see...

Consider the data for confirmed cases (= positive tests) vs day diagnosed for San Diego County.  The plot (available here) shows that after two weeks of cases below 100/day (April 5-April 19), we have suddenly seen a return to numbers over 100 (and the two biggest single day counts of the entire period).  Is this due an increase in the spread of the virus or is there another reason behind the counts increasing?  Well the county has helpfully also included a plot of the number of tests reported by date (available here).  Notice that the number of tests reported is also higher of late.  Someone who wants to see if the case increase is simply due to the test increase merely needs to either 1) plot the number of positive cases/day vs the number of tests reported/day and see if the numbers are correlated or 2) plot the rate of positive cases by day and see if these appear uncorrelated.  The two plots below are what you get for these (note that the data prior to March 21 shows a markedly higher rate of positive cases, much higher than remainder of the data, and a particularly small number of tests, so as to not obscure the relationship I have shown the data from March 21 onward).

  

The plot on the left is the number of positive test vs the total tests, while the plot on the right is the fraction of positive tests by date.

To me the plot on left shows a fairly strong positive correlation and the plot on the right looks like noise but with some non-zero average value.  The slope on the left is 0.623 (note this is not the standard least squares line, but rather is the line you get by a least squares fit where the intercept is forced to the origin - zero tests must yield zero positive results).  The average value for the fraction of positive cases is 0.069 +/- 0.001.  These are in sufficient agreement that I claim that the positive cases reported is strongly influenced by the total number of tests (and will tend to about 1/15th of the total tests).  This means that as the testing increases we will see an increasing number of positive tests.  This is unavoidable!  Hence, we will never meet both criteria.

If it will, in fact, be nearly impossible to see decreasing cases with increasing testing we should ask what will happen if the country won't reopen until we see a decrease (from presumably actually having the virus go through the population).  If the data above is an accurate reflection of reality (vice the far worse possibility that the tests have a false positive rate of 6%), then at any one time we have steadily had about 1/15th of the population infected with active virus.  If this is the stable rate in the locked down economy, and if the infections tend to last 10 days, then it will take something like 90 days to run through enough of the population (~50%)  to really see the cases start to wane, and something like twice that to be starved out of possible hosts.  We simply can not maintain this lock down that long.

I see no reasonable path forward, that is based on actual current data, other than begin a prompt reopening of the majority of the economy as per my post of April 19 and this much better written description of the same thing by someone who actually has credentials.

Aside #1: There was also a statement by that waste of oxygen, also known as the WHO, that it may not be the case that getting the virus and then recovering would provide immunity from future infection.  Well if that's the case, the lock down is a waste of time and we are all truly f****d.  If it is just another case of them trying to make everything look worse, then all the more reason to confine the WHO to the trash heap of history, where it belongs.

Aside #2: Speaking of just "trying to make things worse", I also wonder if the way the CDC keeps redefining how to count cases (and deaths) will also make the cases continue to rise.  The reason I could see the CDC doing this is that they want to make things look as bad as possible (just like the WHO), in order to make their early models less inaccurate.  They need something like that to happen or no one will take them seriously on future predictions.  I hope that isn't the case, but I have to look at all possibilities.

Sunday, April 19, 2020

Reopening the Country - Some Musings on How I Think It Should Be Done and Why.


It is becoming increasingly clear that all over the country people are getting antsy about getting back to something at least approximating their normal life and freedoms.  At the same time, the government has caused us to pay a huge price in those freedoms and in the economy at all levels.  These sacrifices have allowed us to get to where we are now vis a vis the CoViD-19 pandemic (past the peak and seeing general reductions across the board, having never over taxed the hospitals - with the probable exception of a short time in New York City, NYC).  Every time I run my model, the cases, and the deaths jump like crazy when the restrictions are removed.  It seems obvious to me that to simply walk away from them would be irresponsible and foolhardy. 

However, every plan I've heard so far is based on being able to be reactive to the changes that come about as we reopen.  This is unreasonable for two several reasons: 1) we do not have the testing capabilities to determine who really has the disease, or who has had the disease, and 2) we know that the previous outbreak was circulating for an extensive period of time prior to it becoming obvious, thus making any reactive plan doomed to be too little, too late, and 3) doing lots of little steps would just spread the pain out over a longer time.

No, reactive plans are not the right approach.  We need a path where we can confidently calculate the risks and the outcome.  In other words, we need to be able to predict the outcome, not sit back and try to respond to it.  Happily, there are several recent reports that finally give us the data we need to be able to predict with a reasonable safety margin.

First, there have been several recent (and I use the next word loosely) studies that give us a reasonable handle on the true number of cases that have already happened.  There are (to my limited knowledge) at least five results that allow us to determine a number for the ratio of the total cases to the confirmed cases, R_cases.  The first is mentioned in my blog of April 7.  In this event, the entire town of Vo, Italy was tested for active virus.  Some simple math led to the conclusion that R_cases(Vo) ~ 130.  Next came a report from Chicago (see my blog from the 14th).  Here a drive through testing site tested for the presence of antibodies.  They reported that 30-50% of the people tested showed such antibodies.  In this case the math says that R_cases(Chicago) is in the range 110-190.  The third was from the hot bed of American CoViD-19 cases, non-other than NYC.  The R_cases(NYC) comes out to 33 (for details see my blog from the 16th). Finally, over the last couple days there are reports from Santa Clara County, CA and Chelsea, MA (the reports are here and here, respectively).  For these you can easily extract the ratios as R_cases(SCC) ~ 50-80  and R_cases(CMA) ~ 16.  Now there isn't any safe/practical/justifiable way to conflate these into one number, but I think a safe estimate is around 40.  All this implies very strongly that we are much farther on the way to developing herd immunity than anyone would have guessed just a few weeks ago, but how far?

It is quite easy to find data on the web that allows one to compute the relative infection rates for various age groups (in this case see this CDC page).  My calculation says that the infection rate (based on the number of confirmed cases) for those aged 0-18 is less than 0.02%, about 0.15% for those 18-45, 0.22% for those 45-65, 0.18% for 65-75 and 0.26% for those 75 and older (this is for a case load of about 497,000 confirmed cases - which was about April 10).   If we multiply these by 40, the actual infection rates can be estimated at 0.8%, 6%, 8.8%, 7.2%, and 10.4%, and these could easily be anything from a factor of 2 lower, to a factor of 4 higher.  Also the confirmed cases are now more like 700,000, so that's another factor of 1.4.  Nonetheless, we do not yet approach anything like herd immunity (except at the highest of the possible factors, and even then we'd be at something like 50%).

Second, the data now clearly tell us who can, with a relative safety risk, allow themselves to be exposed, and who should not.  All the data point very strongly to the disease being hardest on the elderly [I hate that that world applies to me... I really thought that getting old would take longer...] and those with significant underlying health issues, and is extremely dangerous for those who fall into both categories.  The relative death rates per 100,000 people in the same age groups as above are, less than 1, 11.3, 96.7, 311.7, and 778.8, respectively (from the NYC Health Dept CoViD-19 data webpage).  Based on some CDC numbers (which produce similar rate numbers) I can estimate that the rates for a breakdown of the 45-65 group into 45-55 and 55-65 would be something like 58 and 135. 

So how can we start to reopen as soon as possible?  Reopen slowly, with care, but only for age groups below 55 and only for those without substantial risk factors (see this CDC page, and links thereon).  These groups must do what they can to  keep the spread rate below about half of what it was, (still practice social distancing and extra hygiene practices).  This allows us to get a large fraction of the population back to work with minimal risk.  Even in a worst case (if there is less than 10% of the population currently exposed/immune and the spread rate is very high) the total number of additional deaths would almost certainly remain at or below what we have currently seen and the number of people stressing the health systems would likewise stay below the capacity of the system (the intrinsic death rates for these groups are more than 10 times less than for the most at risk age group, which completely dominates the current statistics, and if those with underlying conditions are removed from the 45-55 age group the death rate is likely to be well below the 58 quoted above).  The big risk in this plan is folks that have been 'reopened' may not keep a hard safe distance between themselves and those who are not reopened (and are much more at risk).  This is a real risk.

But, that's my plan. We would really want to emphasize over and over and over how important it is to not go see those who are still in isolation. We'd want to test as much as possible, and contact trace as well, but it really looks to be a plan that would result in less risk than a general reopening, and would have much less risk of really dire consequences. We'd definitely want to have grocery stores, etc, have special hours for those folks not in the reopened groups (and those folks would probably want to have as much as possible delivered or to arrange to pick it up).  People still won't get to go hug grandma and grandpa until a vaccine has been widely distributed, but such is life.

Thursday, April 16, 2020

CoViD-19: Not Much More Real Data, But Enough to Figure With...

So I'm sure you've seen the report about women checking in to a maternity ward in NYC (story here). While it is always dangerous to give any one bit of data too much credence, this one seems worth following to it's logical conclusion...

First a quick summary: a hospital tested every woman entering the hospital to give birth for CoViD-19. They collected the data over the two weeks (14 days) March 22 - April 4. There were 215 admissions and 33 of them tested positive for the virus. Of them only 4 "had symptoms".

The first question that came to mind was "Had four non-pregnant women called their doctors and described their symptoms, how many would have gotten tested?"
I can't know for sure, but I know from my son (who lives in New York City, NYC) that when he called in with the basic symptoms and a low grade fever, they did not give him a test. So let's guess that maybe 1 would have been tested. In that case, out of the sample of 215, there would have been 1 confirmed case.

The second question, aiming at the same number, would be "Well if we drew 215 random people from NYC, how many would be from among the Confirmed cases?"
We know that NYC reported a total of 62,230 new cases during that time period (see data on the excellent NYC CoViD-19 website). NYC's population is given (Google search) as various numbers from 8.4 million to 8.7 million. Let's use 8.6 million. Further, data suggests that for all but the most severe cases the recovery period is about ten days. So we'd expect to be drawing 215 random people from 8.6 million, of which each day there about 44,450 active confirmed cases (44,450 = 62,230 X 10 / 14). A slightly incorrect but close enough answer is: 44,450 * 215 / 8,600,000 = 1.1 Hence, we'd expect about 1 Confirmed Case.

Those two estimates are not in bad agreement. So let's assume that there really would have been 1.

That leads to another question, "If we see 33 actual cases in the 215, when we'd expect 1 confirmed case, then how many actual cases would we expect there to be in the city?" Well, the ratio of Confirmed to Actual cases in the 215 is 1:33, there would be about 62,230 X 33 ~ 2,000,000 total new cases during that 14 day period. Now extend that to the total time the cases have been running through the city. That means we take the total number of confirmed cases, which is 117,565 (and still climbing!) and do the same multiplication. That works out to something like 4 million. That's nearly half the population!

Further, there is data from Vo, Italy and some hints from Chicago that the ratio of total cases to confirmed cases is in the range 100:1 to 200:1 (see the previous two posts). If you believe those (and the data from the pregnant women don't really exclude these given the small numbers of them tested), then we are talking numbers so big that essentially everyone in NYC has had CoViD-19. If true, we can expect the number of new cases there to absolutely crash over the next week to 10 days. It will be interesting to see. On the other hand, if the numbers don't crash, then I'll have to toss my existing models out and start again with a reduced ratio like 33.






Tuesday, April 14, 2020

CoViD-19: A Follow Up

Words of Caution:


See the post below.  They all still stand!

Background:


Last week I presented some results from a quick and dirty model, utilizing the best parameters I could find, and the simplest assumptions that allowed me to fit the data.  Well it's been an interesting week, to say the least.  I have had to make a few changes to fit the new data.  So without any further ado...

Data:


I am still using the same data sources for cases and deaths as I did last week.  I just can't find anything that even remotely seems better.  I now have enough CDC death numbers that I'm showing those as well.

Results:


There are four plots below that graphically show the fits.  I am satisfied that while they are some what ad hoc, they agree to a more than adequate extent.

A                                                                              B
Figures 1A and B. Two plots of the Worldometer and CDC Data and the Model results (the one on right, B, being a blow up of the y-axis to show the quality of the agreement at large numbers).


Figure 2. Plots of the Day-to-Day cases numbers.  Note that the data are extremely noisy over the last 10 or so days.


Figure 3. Data for the number of deaths.

It is clear that the model is doing a good job of capturing the dynamics of the CoViD-19 outbreak.since last week.

The conclusions I would draw have changed only somewhat taking into account that last week the model was predicting far fewer cases, because I was completely taken in by the wiggle you can see in the day-to-day case data from around 3/27-3/31.  Beyond that the comparison to last weeks "suggestions" are:
  1. The stay-at-home orders and the social distancing being practiced across the country are having a clear and downward pressure on the cases and will also show a reduction in the deaths as well.  This still holds.
  2. There is some indication that there is some small positive influence on disease outcome (people not dying!) from whatever treatments are being provided. This still holds.
  3. The number of infectious cases is still very high and these practices must be maintained for the foreseeable future if we are to clear the virus from our population. This still holds.
  4. The maximum number of new cases (i.e., the number of people who get infected on a specific day, and hence are now infectious) is likely behind us (model says it occurred on or about March 28). The conclusion holds, but the date at which it occurred was April 2.
  5. The maximum number of infectious people on any given day is also likely to be behind us (model says this happened on or about April 2 or 3). The conclusion holds, but the date at which it occurred was April 7.
  6. The maximum number of newly Confirmed Cases on any given day is likely to have happened, be happening, or about to happen (model shows it somewhere in the span April 5, 6, or 7).  The conclusion holds, but the date at which it occurred was April 10.
  7. The maximum number of new deaths is likely to still be in the future (somewhere around April 8 - 11).  The conclusion holds, but the date at which it will occur is on or about April 16.
The other major change is that the length of time to fully clear the infection is MUCH longer.  I won't even say how far out the model predicts them, other than to say, some of them extend into 2021.  I should also mention that the predicted number of deaths has jumped to about 63,000.  Also the number of people actually infected (and almost entirely recovered) ends up being just under 150 million of the total population of 330 million.

I should also mention that two of the primary factors that influence the model are the number of total cases versus the number of confirmed cases.  These results are based on the same numbers I used last week.  Recently KT has pointed me to two rather interesting reports (Thanks!).  The first is a report that claims the mortality for CoViD-19 is around 0.06% (see the second bullet under the New Studies under the April 12, 2020 heading here).  That is in quite good agreement with the parameter I have been using (7.6%/137. ~ 0.056%) The second is a rather interesting report (even though somewhat anecdotal) from Chicago that 30 - 50% of the people tested shows signs of having had CoViD-19.  If the higher figure holds that suggests that the number of hidden cases might be significantly higher than the assumed parameter of 137 (perhaps as high as 190, but also as low as 110).  A few runs using the high end parameter resulted in a nearly identical looking fit, with essentially the same dates for the peak, but with slightly shorter time to fully clear the cases (but still out to December), around (but below) 60,000 deaths, and something like 190 million total cases.  .The lower value will push the numbers the other way.

Tuesday, April 7, 2020

CoViD-19 – Some “In-the-Box” Thinking


Words of Caution:

First let say the following things clearly and for the record:
  1. I am not an epidemiologist although I have had some familiarity with chemical kinetics equations, which are very similar to the epidemiology equations for disease spread,
  2. I do not warranty any of what follows,
  3. I strongly urge everyone to continue to follow the guidance of the US CDC and all guidelines and orders of Local, State, and Federal authorities, and
  4. I to specifically urge that you NOT use what appears here to make personal decisions.

Background:


During this time while I am stuck at home (i.e., In-the Box), I decided to occupy some of my time by trying my hand at modeling the current Coronavirus (CoViD-19) outbreak here in the US. After finding a free FORTRAN compiler, I started programming the equations that describe the standard epidemiology equations for disease spread. For simplicity I chose to use the discrete version rather than the differential forms. I have described the Model (both the equations and how I chose various parameters) at the end of this post, for those who care about the details.

Data:


The data used here are those available on two different websites for the number of CoViD-19 Cases and Deaths for the United States. The data labeled Worldometer Data is from the website here and the data labeled CDC Data is from the website here. The Worldometer website has interactive plots of both the cases and deaths, and the data I have here was extracted from them by hovering over the various points and recording the data. The CDC website has the case data available in a table that can be scrolled, but I have seen no historical data for the deaths, and as I have been recording those data only since 5 days ago I have not shown them on the plot, nor adjusted any of the parameters to try to match them.

Results:


Below are the current (April 6, 2020) plots for the available data and the model fits. I am satisfied that these are in sufficient agreement with the past that I am willing to say a few more things.

Confirmed Cases vs. Days since the 100th Confirmed Case for the US.

Deaths vs. Days since the 100th Confirmed Case for the US.

This model suggests that: 
  1. The stay-at-home orders and the social distancing being practiced across the country are having a clear and downward pressure on the cases and will also show a reduction in the deaths as well.
  2. There is some indication that there is some small positive influence on disease outcome (people not dying!) from whatever treatments are being provided.
  3. The number of infectious cases is still very high and these practices must be maintained for the foreseeable future if we are to clear the virus from our population.
  4. The maximum number of new cases (i.e., the number of people who get infected on a specific day, and hence are now infectious) is likely behind us (model says it occurred on or about March 28).
  5. The maximum number of infectious people on any given day is also likely to be behind us (model says this happened on or about April 2 or 3).
  6. The maximum number of newly Confirmed Cases on any given day is likely to have happened, be happening, or about to happen (model shows it somewhere in the span April 5, 6, or 7).
  7. The maximum number of new deaths is likely to still be in the future (somewhere around April 8 - 11).
  8. The last new case is likely to be infected on or about May 22nd (!).
  9. The last death may occur on or about the same day and will represent something like the 24,000th death.
  10. The last infectious person will clear quarantine some days after that (10-24 days are typical recovery times I've seen quoted, so that works out to June 1-15).
The bottom line I take from the model is that this epidemic is no where near done, although we may start seeing some positive news soon, possibly this week. Also, when things start to look really good we must all take a deep breath and continue the stay-at-home and social distancing for at least 2 weeks more than we will want - or seem to need. This will be critical in assuring that the last few cases find no new victims.

An additional result is that the model predicts about 82 million people will have had the disease and recovered. This will be interesting to compare to any retrospective antibody testing that are done on the population of the US.

Finally, I suspect that the model will have to be adjusted to reduce the drawdown factors for both cases and deaths, which will likely move all the dates which are in the future farther out (and will result in higher figures for both the number of people that recover and, unfortunately, the number that die. On the other hand, if these turn out (miraculously) to be correct, remember that you heard it here first.

The Model:

The variables:


J = day number
U(J) = number of uninfected people on day J
N(J) = number of newly infected people on day J
I(J) = number of infectious people on day J
C(J) = the number of people who are confirmed to have the disease on day J
TC(J) = the total number of confirmed cases through day J
D(J) = the number of people that die on day J
TD(J) = the total number of deaths through day J
AR(J) = the number of asymptomatic cases that recover on day J
CR(J) = the number of confirmed (=symptomatic) cases that recover on day J
R(J) = the total number of people who recover on day J
TR(J) = the total number of people who have recovered through day J

Assumptions:

  1. There are cases of CoViD-19 which can be either symptomatic or asymptomatic. For the purposes of simplicity the symptomatic people will all be assumed to be confirmed, and the asymptomatic cases will be assumed to not be confirmed.
  2. People who are asymptomatic are assumed to be infectious for 10 days(they recover on the 11th day).
  3. People who are symptomatic (and hence, confirmed) are assumed to be infectious for only 7 days as they are assumed to become confirmed on day 7.  All confirmed cases are assumed to either be in the hospital or in quarantine and, therefore, in both cases it is assumed that the conditions do not allow them to infect anyone.
  4. Confirmed cases all end up either recovering on the 24th day from their initial infection or dying on the 11th day after infection  (originally I assumed this would be around day 15, but to get a reasonable fit I had to change it). [Note: These lag times have no impact on the model  beyond which day they show up in the totals, as they were removed from the infectious pool when they were confirmed.]

The Equations: 

[Note: the differential equations are available by searching for the SIR model]

The Base Model



J = J + 1
N(i) = k1 * U(i-1) * I(i-1)
U(i) = U(i-1) – N(i)
C(i+7) = k2 * N(i)
TC(i) = TC(i-1) + C(i)
D(i+10) = k3 * C(i)
TD(i) = TD(i-1) + D(i)
AR(i) = N(i-10) – C(i-3)
I(i) = I(i-1) + N(i) – AR(i) – C(i)
CR(i) = C(i-24) – D(i)
R(i) = AR(i) + CR(i)
TR(i) = TR(i-1) + R(i)

Note 1: all these equations use integers as inputs and outputs. The math is done as real numbers and any non-integer part is included or excluded, as an extra person, based on a random number.
Note 2: In fact the calculation of N(i) is the sum of ten terms of this form, which actually use N(i-10) through N( i-1) and C(i-3) through C(i-1), vice I(i-1), as will become clear in the next section.

This base model basically results in essentially everyone in the model getting infected. So for the US total population of 330 million the death total comes out at slightly less than 200,000 deaths. This however does NOT fit the data, after about the first couple weeks. It just goes on up exponentially until U goes to essentially zero (actually 69).

Input Parameters/Information:


The initial condition is that U(0) = 330,000,000 and all others are 0. Then N(1) is set to 1, and off the model goes.

k1 is not just one number, but rather it is actually an array of ten multipliers based on which day it is from the the day a person gets infected. The array varies such that the actual infectious-ness of an individual is such that: for each of the 3 people they infect, they will infect (on average) 1/2 person total over days 1-3, 1/2 person over days 8-10, and the remaining 2 people are infected over days 4-7 (the numbers being the number days after they were infected). This is a crude approximation of the data given under the heading “'Characteristic' Infection progression in a single patient” here.

k2 = 90. / 3300. This based on the only study I have heard of that may properly define this ratio. This was data reported for the city of Vo, Italy. It was reported in the Wall Street Journal, WSJ. I saw it where that article was being discussed here). Hat tip: KT

k3 = 8% From the same WSJ article via the same blog post. Again Hat Tip: KT

Adjustments to the Base Model

These adjustment have been implemented to be able to match the actual number of cases and actual number of deaths over the current data set.

Adjustment 1: As can be seen in the plot for the number of cases, the initial exponential growth (the straight line in this log plot) begins to roll over on or about the 17th day (this is not clear in the graphs shown but if you look closely, on a greatly expanded scale, you can make it out). This date works out to March 20. This would seem to match fairly well with the onset of stay at home orders (for example NY Governor Cuomo on March 20 and CA Governor Newsom on March 19). The actual adjustment that gives an curvature rate that matches the data is to reduce the daily infection rate by a factor of 0.885^(i-16).

Adjustment 2: The actual death rate did not match using the crude figure of 8%. It required a reduction to 7.6%. This can easily be accounted for by any of several factors including: better more complete testing in the US vice what Italy had accomplished, a healthier population (Italy has a significantly higher rate of smoking), or some other minor factor (like simple rounding in the value of 8%).

Adjustment 3: Much as the number of cases was adjusted it became clear that matching the number of deaths requires an additional similar factor. The best match (by eyeball) uses it starting on day 23 (real world = March 26) and a factor of 0.96^(i-22). It is not clear what might case this, but I suspect that it is a reflection of some actual increase in the ability to treat patients.

Addendum:


This post, and the data and model results it is based on are through April 06 (actual data end April 05). Today, April 07, the data from both Worldometer and CDC show a significant spike upwards in both cases and deaths on April 06. The cause for this is unclear. I will not make any additional model changes for several more days.