Count what actually arrived at your practice yesterday.
A dozen phone calls, of which perhaps three mattered. Four web inquiries, of which one was a real potential client and three were vendors, SEO solicitations, or somebody in another state. A message in the client portal. Three e-service notices from the court. Some mail. A handful of emails from opposing counsel that were genuinely urgent, mixed into a hundred that were not.
Now count the information in all of that. It is not much — maybe eight facts you needed to know and four decisions you needed to make.
That gap is the entire problem. The volume of inbound is trivial. The volume of interruptions generated by inbound is enormous, and interruptions are paid for in the only currency a practice cannot buy more of.
I have written elsewhere in this section about the mental load of running several pieces of work at once. Inbound is the same problem arriving from outside. The difference is that you did not choose it and cannot schedule it — which makes it worse, and makes it the higher-leverage thing to fix.
What a message actually costs
Be precise about the cost, because “it’s just a quick call” is how the whole day disappears.
A call that takes four minutes does not cost four minutes. It costs four minutes, plus the reconstruction of whatever you were holding when the phone rang, plus the residue — the low-grade background process that keeps chewing on the call for the next twenty minutes whether you want it to or not. If you were deep in a brief, the real cost is closer to half an hour.
Now the crucial asymmetry: that cost is identical whether the call mattered or not. A four-minute call from a client with a genuine emergency and a four-minute call from somebody selling you a search-engine package cost exactly the same amount of attention. The second one just returns nothing for it.
Which means the highest-value thing any inbound system does is not summarizing, routing, or drafting responses. It is making sure the second call never reaches you at all.
Classify at the edge
The design principle is to classify as early as possible — at the point of arrival, before anything reaches a human — and to route by consequence rather than by channel.
Channel-based organization is what most practices have by default: phone messages here, email there, portal messages in a third place, court notices in a fourth. It feels organized. It is actually the worst arrangement available, because it groups things by the accident of how they arrived rather than by what they demand from you. An urgent client call and a robocall are both “phone.” A dispositive order and a newsletter are both “email.”
Consequence-based routing groups by what you have to do about it. Our inbound splits into four:
Needs a lawyer today. A client in trouble, opposing counsel on a live deadline, a court. Goes to a channel that is allowed to interrupt.
Needs a lawyer, not today. Substantive but not urgent. Goes to a channel you read at a chosen time — which is a very different thing from a channel that reads you.
Needs somebody, not a lawyer. Scheduling, records requests, billing questions, routine intake.
Needs nobody. Solicitations, wrong numbers, hang-ups, spam. Logged, never surfaced.
The fourth category is the one that produces the gains, and it is the one most systems refuse to implement, because deleting things feels reckless. It is not reckless if it is logged and searchable. It is only reckless if it is gone.
Separate channels, deliberately
One practical decision that matters more than it sounds: put each category in its own channel, and let the channel’s identity carry the urgency.
Ours are separate feeds — one for calls that matter, one for anything from opposing counsel, one for new leads, one for portal notices, one for new court filings. A glance at which feed lit up tells you the shape of the thing before you have read a word.
The point is not tidiness. It is that the notification itself becomes informative. When everything lands in one inbox, every alert requires reading to triage, so every alert is an interruption. When feeds are separated by consequence, most alerts can be deferred without being read, because you know what kind of thing it is from where it landed. That is the whole trick: not fewer messages, but fewer messages that must be read now to find out whether they must be read now.
A corollary: the urgent channel must stay genuinely rare. The moment it starts carrying ordinary traffic, you stop trusting it, and once you stop trusting it you read everything again and you are back where you started. Guard it aggressively. One or two things a day, or the channel is broken.
Suppression needs an audit trail
The obvious objection: what if the classifier is wrong and suppresses something that mattered?
It will be, occasionally. The design answer is not to stop suppressing. It is to make suppression reviewable rather than invisible.
Three rules:
Log everything, surface selectively. Nothing is deleted. A suppressed call has a record with its transcript or summary, its classification, and why. Searchable forever.
Make “show me what you filtered” a first-class command. You should be able to ask what got suppressed today and read the list in ten seconds. In practice you will do this weekly at first, then monthly, then almost never — because the answer is boring, which is how you learn to trust it. That trust is the actual product.
Tag rather than suppress when a category is uncertain. For a class of message where a miss would be costly — an inquiry that might be a real client — mark it and let it through. Suppression is for things you are confident about. Spam that looks like a lead gets tagged as suspected spam and surfaced anyway; the tag does the work without the risk.
Alert-only is a legitimate design
A constraint worth generalizing from, because it comes up constantly with practice management software.
Some systems will not let you read message content programmatically. The client portal is a common example — you can learn that a message arrived, from whom and when, but the API will not give you the text. The instinct is to treat that as a blocker and abandon the integration.
It is not a blocker. Knowing that a message exists is most of the value. The failure mode you are actually trying to prevent is a client message sitting unread in a portal for four days because nobody logged in. An alert saying “a portal message arrived on this matter” prevents that completely. You then open the portal and read it there, which takes fifteen seconds and is exactly what you would have done anyway.
The general form: when a system will not give you content, take the signal. A notification that something happened, routed to the right feed, solves the “nobody noticed” problem — and “nobody noticed” is the problem that actually hurts practices. Do not let the unavailability of the ideal integration cost you the useful one.
One caution: verify the limitation rather than assuming it. We probed the API before concluding message text was unavailable. Building an alert-only integration because you did not check is a different thing from building one because you did.
Court filings belong in this system too
Docket monitoring is usually thought of as a separate category, and it should not be. It is inbound.
If your jurisdiction uses electronic service, every filing in every one of your cases already arrives as a notice with the document attached. That is a complete new-filings feed, delivered to you, at no cost — and in most practices it sits mixed into an email inbox where it is indistinguishable from everything else.
Pull those notices out, match each to a matter, download the document, and present the day’s filings as a single list. That is a docket alarm, built from something you are already receiving.
The reason to treat it as part of the same inbound system rather than a separate tool: it is the same problem — a low-volume, high-consequence stream that must not be missed and must not interrupt. It wants exactly the same treatment. Its own feed, matched to matters, read once a day.
What to build first
If you are starting from nothing, the order matters, because the early wins are what buy you the patience for the rest.
1. Suppression. Before any summarizing, classification, or routing — just stop the noise. Filtering out solicitations and hang-ups from the phone log is the single largest attention return available and it is nearly trivial to implement.
2. Separate the urgent channel. One feed that is allowed to interrupt, with a deliberately narrow definition. Everything else goes to a default feed you read at chosen times.
3. Then classify the rest. Split the default feed by consequence as patterns emerge. Let the actual traffic tell you what the categories are rather than designing them up front.
4. Then add the low-volume streams. Court filings, portal notices, web inquiries. These are easy once the routing exists and they are not where the pain is.
Resist the temptation to start at step four because it is the most technically interesting. The gain is at step one, and it is available in an afternoon.
The point of all of this
None of this makes anyone a better lawyer, and I would not claim otherwise.
What it does is protect the hours in which you can be one. A practice in which every inbound message costs half an hour of scattered attention is a practice that produces less work, of lower quality, at higher cost — and passes that cost to clients, or to the people who never become clients because your rate had to cover it.
That is the connection back to everything else in this section. The number of cases a lawyer can carry well is bounded by attention, not by hours in the day. Anything that stops attention from leaking raises that ceiling. And a higher ceiling is, eventually, a lower price.
General commentary on practice management. Not legal advice. No client, caller, or matter is described in this article; all examples are invented.