Exploratory analysis of 1,898 NYC food-delivery orders, turned into decisions a business can act on: which restaurants earn the promo, where delivery lags, and where the revenue really comes from.
A NYC food-delivery platform has 1,898 orders across 178 restaurants sitting in a table. The analysis isn't interesting until someone has to act on it — so the brief was framed as three decisions the business actually faces: which restaurants earn a promotion, where delivery is underperforming, and where the revenue really comes from.
That framing changes what counts as a finding. A histogram of order costs isn't a result; "71% of orders clear the higher commission tier" is. Every chart in the notebook had to end at a decision or it didn't earn its place.
Standard exploratory analysis, run in the standard order — but pointed at the three questions rather than at the columns:
Built with pandas, NumPy, Matplotlib and Seaborn. No modeling — this one is entirely about reading the data honestly.
Demand is concentrated. Shake Shack alone accounts for roughly 11.5% of all orders, followed by The Meatball Shop, Blue Ribbon Sushi, Blue Ribbon Fried Chicken and Parm. American and Japanese cuisines lead on weekdays and weekends alike. That top-five list is the promotion shortlist, and it's stable enough across the week to act on.
Delivery has a weekday problem. Mean delivery runs about 24 minutes overall, but that average hides a consistent split: 28 minutes on weekdays against 22 on weekends. Lower-rated orders also correlate with longer prep and delivery times, which points operations attention at a specific window rather than at delivery in general.
Revenue sits in the big tickets. With commission at 25% above $20 and 15% below, and 71% of orders clearing $20, the platform earns roughly $6,167 across this order set. The cost distribution is right-skewed in a way that suggests two distinct order types — individual and group — worth targeting separately.
The most important finding was a hole in the data. 39% of orders — 736 of them — carry no rating at all. That's the result I'd lead with in a real meeting, because it invalidates the obvious next move. Any conclusion drawn from the rated subset is a conclusion about the kind of customer who leaves ratings, not about customers. Before this business optimizes on ratings, it has to fix collection.
Averages hide the thing you're looking for. The overall 24-minute delivery time is unremarkable and would have ended the analysis. The 28-versus-22 split underneath it is the actionable version of the same number — a reminder that in EDA the aggregate is where you start, not what you report.
The full notebook — every distribution, the cuisine and timing breakdowns, and the commission calculation — is on GitHub.