How to choose a custom software development company by what its first month shows
Every agency on your shortlist promises skill, quality and good communication. None of that can be checked on a sales call. What a team hands you in its first four weeks can, and that is the test worth running.
How do you pick a custom software development company?#
Pick the custom software development company whose first month you can verify: working software shown every two weeks, code in a repository you own, and a signed copyright assignment. Because each of the three sits on the record, you can check it yourself rather than take a pitch on trust.
This idea comes from public guidance. The U.S. government's 18F team wrote its De-risking Government Technology guide for agencies that hire outside teams. The General Services Administration announced its revised edition on 11 September 2024. The guide says quality indicators "should be reviewed at every sprint, usually every two weeks". Then it sets the standard that "at the end of each sprint, all code is delivered to a government-owned repository".
The third part is legal. The U.S. Copyright Office published Circular 30, revised in August 2024. It says a commissioned work counts as made for hire only when it meets four tests, and one reads: "There must be a written agreement between the party that ordered or commissioned the" work and its maker. In short, payment alone does not settle ownership.
So the question you ask of a shortlist changes. Instead of asking who sounds most experienced, ask what each firm will let you check in four weeks.
| What to check | How you check it | Who sets the standard |
|---|---|---|
| Working software | Watch a demo at the end of every two-week sprint | GSA 18F guide, September 2024 |
| Code you own | Open your own repository after each sprint | GSA 18F guide, September 2024 |
| The copyright | Read the signed written agreement that assigns it | U.S. Copyright Office, Circular 30, August 2024 |
When is a development company the wrong fit?#
If many other organisations need the same thing, buy a product, after the build vs buy check in the custom enterprise software development process guide; if the software is your permanent core, hire in-house; if you only lack hands, add contractors. A development company fits between those three: work that is specific to you, needed now, and too big for one or two extra people.
18F puts the first rule plainly in its 2024 software solutions chapter. Off-the-shelf software "is the right choice for meeting a need that" many others share, such as email. But building your own email system would waste time and money.
The trap is the middle path. Buying a product and then reshaping it until it no longer takes the vendor's updates is what 18F calls "unrecognizably modified off-the-shelf" software. The guide says it "increases the risk of project failure" and usually locks you into one vendor. In effect, a heavily customised product is custom software with extra rent.
The public data points the same way. In a study of 1,355 public sector IT projects, Budzier and Flyvbjerg found standard software projects ran into large overruns more often than bespoke ones.
Show data table
| Item | Value |
|---|---|
| Standard software | 26 |
| Bespoke software | 18 |
| IT architecture | 16 |
| IT infrastructure | 14 |
Standard software overran by more than 25% more often than bespoke software, 26% against 18%.
The other two exits are simpler. If the software will run your business for the next ten years, the people who understand it should work for you. Our nearshore, offshore and in-house guide weighs that choice in depth. Finally, if you have a team and a plan but too few hands, one or two contractors are easier to manage than an outside team.
How often do software projects go badly wrong?#
In a study of 1,471 IT projects the average cost overrun was 27%, but one project in six overran by 200% on average. Those are the figures Bent Flyvbjerg and Alexander Budzier published in Harvard Business Review in September 2011.
27%
Average cost overrun across 1,471 IT projects
200%
Average cost overrun of the one project in six that was a black swan
One project in six overran by 200% on average, far beyond the 27% average.
| Option | average cost overrun, percent of budget |
|---|---|
| Average cost overrun across 1,471 IT projects | 27% |
| Average cost overrun of the one project in six that was a black swan | 200% |
Source: Flyvbjerg and Budzier, Harvard Business Review, September 2011
A 27% overrun hurts, but most companies in that 2011 sample would survive it. Yet a project that costs three times its budget is a different event. In the authors' words, "Fully one in six of the projects" was a black swan. Those projects overran cost by 200% on average and ran almost 70% late.
Later work found the same shape. In October 2022, Flyvbjerg, Budzier and four co-authors studied 5,392 IT projects completed between 2002 and 2014 and found the overruns follow a fat-tailed power law.
Meanwhile, 18F's guide opens with the line "Only 13 percent of large government technology projects succeed."
Therefore the danger to plan for is not a modest overrun. Instead, it is the rare project that runs for a year before anyone sees that it is lost. Choosing a vendor is mostly a way to make sure you would see that early.
What should you settle before you talk to any vendor?#
Settle four things before the first call: the problem in your users' words, who makes product decisions, what success means in numbers, and what you will not build. Because a vendor can only be judged against decisions you have already made, these come first.
First, the problem. Write it the way the people who will use the software would say it, not as a feature list. Second, the decider. 18F's principles tell agencies to "Identify and empower a full-time, in-house product owner to lead the project". Since that person reviews the vendor's work every sprint, they need the time and the authority.
Third, the measure. 18F gives examples such as "a 20 percent increase in digital application submissions" or a feature shipped every two weeks instead of once a quarter. Finally, the limits: what is out of scope this year.
The UK's audit office explains why this cannot be left to the supplier. Its report on technology suppliers, published 16 January 2025, says design and architecture in digital programmes "need to be done to a sufficient extent by the department in advance of contracting". In its own words, "A larger percentage of digital programmes were rated" red than any other type of programme in 2022-23.
Show data table
| Item | Value |
|---|---|
| Digital (31 programmes) | 16 |
| Infrastructure and construction (73) | 11 |
| Military capability (42) | 10 |
| Transformation and service delivery (86) | 7 |
| All programmes (232) | 10 |
Digital programmes were rated red at 16%, above the 10% for all programmes.
If those four answers are still vague, settle them first. A short product discovery phase is one way to do that before any build contract. Also, the cost guide covers the budget half of the brief.
Which questions separate one vendor from another?#
Ask questions a vendor can only answer with evidence: show us code like ours, name who will write ours, and describe what we will see working in four weeks. But a question answered with an adjective tells you nothing, because every firm uses the same adjectives.
Here are six questions, with what a good answer looks like.
- Can we see code like ours? 18F's buying guide says to "Ask for links to two or three source code repositories" of similar size. A link is a good answer; a PDF is a weak one.
- Who will write our code? Ask for the leads by name, then interview them. Since a firm cannot "lock down all key personnel months before" work starts, ask for two or three names, not ten.
- What will we see working at week two and week four? A good answer names a screen or a workflow, not a document.
- Where does the code live? In your own account, from the first sprint.
- How will we know the code is tested? A good vendor shows its test reports with no extra work.
- What if we stop after one month? You keep the code, the designs and the notes.
| Question | A good answer | A weak answer |
|---|---|---|
| Can we see code like ours? | A link to two or three repositories | A PDF |
| Who will write our code? | Two or three leads by name, ready to interview | Ten names, or none |
| What will we see working at week two and week four? | A screen or a workflow | A document |
| Where does the code live? | In your own account, from the first sprint | In the vendor account until the final invoice |
| How will we know the code is tested? | Test reports shown with no extra work | An adjective |
| What if we stop after one month? | You keep the code, the designs and the notes | Nothing you can take with you |
Once proposals arrive, comparing and scoring them is a separate job. The RFP template covers it.
Which red flags should end the conversation?#
End the conversation when a vendor wants to set its own measures of success, keeps the code on its own servers, or promises one big delivery at the end. Each one hides progress from you until it is too late to change course.
The vendor wants to set its own measures of success.
Your measures come from your brief, not from its proposal.
The code sits on the vendor servers.
You cannot check the work, and you cannot leave.
One big delivery at the end.
Any problem surfaces when the budget is already spent.
A heavily modified product, explained vaguely.
Ask which parts are configuration and which are new code.
The first flag is a vendor grading itself. The 2024 18F guide is blunt about it: "If you allow the vendor to define its own measures of success", you give up one of your strongest checks on quality. So your measures come from your brief, not from their proposal.
The second flag is code you cannot see. When the repository sits in the vendor's account until the final invoice, you cannot check the work, and you cannot leave. Because of that, a vendor who resists your own repository is telling you how the exit will go.
The third flag is a single delivery at the end. For example, a plan with one release at month twelve gives you nothing to judge for a year. As a result, any problem surfaces when the budget is already spent.
There is a quieter fourth flag. 18F notes that some vendors sell a heavily modified product "with inaccurate or incomplete explanations" of what the product can actually do. So ask which parts are configuration and which are new code.
Who owns the code when the work is delivered?#
Paying for software does not by itself make you its owner: under US law a contractor keeps the copyright unless a signed written agreement transfers it. This section is general information about US law, not legal advice, so have a lawyer review your contract.
The law firm Orrick, in a note dated 13 September 2023, puts it simply. U.S. copyright law "assumes the contractor owns all intellectual property rights in the developed software unless the contractor assigns rights to the company in writing". For employees the rule runs the other way, which is why hiring in-house and hiring a firm differ here.
Two details matter in the contract. First, the wording should be a present assignment, such as "hereby assigns", not a promise to assign later. Orrick explains that a promise to assign does not itself move the rights. Second, a "work made for hire" clause helps but is not enough alone. Circular 30 lists four conditions a commissioned work must meet, including the written agreement, and much software falls outside its nine categories.
Also, location changes the rules. The same Orrick note covers France, Germany and the UK, and each one has its own formalities. If your vendor's people work in another country, ask for clauses written for that country.
The practical half is the repository. Code that arrives in your account every sprint is code you hold, whatever happens to the relationship later. And if you already own code from a past vendor, maintenance and support is the route for keeping it running.
What should the first month of work show you?#
By the end of the first month you should have seen two working demos, code in your repository after each sprint, and automated test results you can read. When any of the three is missing, that is the earliest warning you will get.
18F's vendor management chapter calls demos "the only way the government can be truly confident" that a vendor's work is on track. The same chapter lists what a modern team can show at each review:
- A demo of working software at the end of every sprint, even if users will not see it yet.
- Code in your repository, complete and documented, after every sprint.
- Test reports. 18F's own standard is 90 percent code coverage from automated tests.
- A single-command deploy to a test environment, written up in plain language.
- Security and accessibility checks that run as part of each sprint review.
- Week 2
Sprint one ends
A demo of working software, code in your repository, and test results you can read.
- Week 4
Sprint two ends
A second working demo, more code in your repository, and the coverage number. 18F sets its own standard at 90 percent.
- Any time
What you can check yourself
Watch the demo, open the repository, and look for small, frequent commits.
Still, none of this needs a technical background to check. For instance, you can ask for the coverage number, watch the demo, and open the repository yourself. In particular, look for small, frequent commits, which show steady work.
A vendor who cannot show these in four weeks will not start showing them in month six. Instead, the gap tends to grow while the bill grows with it.
How does this play out for a 40-person distributor?#
A 40-person distributor comparing a single month-12 release with two-week sprints would see two working demos from one vendor and none from the other by week four. The example below is illustrative, built from the method above.
Say the distributor needs a tool for dispatch and delivery tracking. First, it checks the products on the market, because many firms need dispatch. It decides to build only because its own routing rules are what win it customers. After that, it talks to two vendors with the same budget.
Vendor A proposes a full specification, a build, and one release at month twelve. Vendor B proposes two-week sprints, a demo after each, and code in the distributor's own repository from day one. By week four, vendor B has shown two working pieces, perhaps a driver's job list and a live map. Vendor A has shown a document.
Now count the blind time. A twelve-month plan is about 52 weeks. So with vendor A, the first demo comes after all 52 weeks, and nothing working is seen until the budget is spent. With vendor B, the first demo arrives at week two, about 4% of the way in.
How much of the plan passes before anything works?
Set the planned length and the week of the first working demo. Illustrative worked example, not a forecast.
Share of the plan passed before anything working is seen
4%
- Planned length in weeks
- 52
A twelve-month plan with a first demo at week two gives about 4%; a first demo at week 52 gives the whole plan.
Time adds risk as well. In the 1,355-project study published in December 2012, every additional year of project duration added 4.2 percentage points of average cost risk. So if vendor A's plan slips into a second year, the distributor carries that extra risk too.
What can the failure data not tell you?#
The overrun studies cover large and mostly public IT projects recorded more than a decade ago, so they show the shape of the risk, not your odds. So a small private build may behave quite differently.
The same 1,355-project study from 2012 shows the shape clearly. In fact, the typical project in it had no cost overrun at all. However, the chance of each bad outcome was real.
| Outcome | Chance | Extra cost, median | Extra time, median |
|---|---|---|---|
| Overruns the budget | 28% | 47% | 38% |
| Black swan | 18% | 130% | 41% |
The 130% here and the 200% from 2011 are both true. They come from different samples and different measures. This study counts any overrun above 25% as a black swan and reports the median one, at 130%. The 2011 study reports the average of its black swans, at 200%. A few huge overruns pull an average far above the median. Both studies put about one project in six in that group.
Read those numbers as a reason to look early, not as a forecast. The projects were large: on average they spent $130 million and ran 35 months. Your project is probably smaller and shorter, and shorter projects carried less risk in the same data.
The method has limits too. Demos, a shared repository and a signed assignment lower the risk; they do not remove it. For instance, a vendor can demo a screen that hides weak code, which is why the test reports and the code itself matter as much as the demo. Also, nothing here replaces a lawyer's review of your contract.
Where should you go from here?#
Compare written proposals with the RFP template, weigh team location in the in-house guide, and read the cost guide before any vendor quotes you a number. Pick the page that matches the decision in front of you.
- Comparing proposals: the software development RFP template shows how to score them.
- Choosing a team model: nearshore vs offshore vs in-house weighs a local team, a team abroad and your own hires.
- Setting a budget: how much custom software development costs explains what drives the number.
- Replacing an old system: how to choose a legacy modernization partner covers that case.
- If you are still asking whether custom software is the answer, start with what custom software development is.
Our custom software development page shows how one team runs its sprints. The product discovery page covers the work before a build. Either way, the checks above work the same with any vendor. A first month you can verify is the whole test, and you can run it without us.