The arrangement is everywhere because it is genuinely useful. A form collects the records, a linked spreadsheet holds them, a few formulas summarise them, and a handful of people read the result. It costs nothing, it took twenty minutes to build, and for a surprising amount of work it is the correct answer. Then something snags. A record needs correcting and the formula column stops lining up. Two entries turn out to be the same person. A summary sheet takes nine seconds to recalculate. Somebody asks a question the table cannot answer.
The useful thing to know is which of those snags are fixable inside the arrangement and which are properties of it. They are different problems, and treating a property as a bug is how teams spend three months rebuilding something that was never going to work.
What the arrangement actually is
Named accurately, a form writing to a spreadsheet is an append only log with a header row acting as a loose schema. That description is not a criticism. It is a real and respectable data structure, and it has three properties that make it stronger than the shared spreadsheet it usually replaces.
Writes do not collide. Two people submitting at the same moment produce two rows. Two people editing the same cell in a shared sheet produce one silent loss. Routing all writes through a form eliminates the most common cause of quiet data corruption in a small team.
The shape is enforced at the point of entry. Every row has the same columns in the same order, because the form decided what could be typed. Data validation rules on short answer questions, such as a maximum character count, tighten this further. A shared sheet has no such guarantee and drifts within a fortnight.
Nothing is destroyed. New facts arrive as new rows, and the old rows are still there. That makes the history reconstructable, which is the single most valuable property for anything that might be questioned later.
Where the arrangement breaks down is not capacity and not speed. It is that a log is not a database, and four things a database provides are simply absent.
The four missing properties
Identity. A database has keys. This log has none. The same person applying to three programmes is three unrelated rows, and the only thing connecting them is an email address typed in a text field, which is not reliable. Forms can collect email addresses either as verified, taken from the signed in Google Account, or as responder input, which is whatever the person typed. Responder input means the joining key can contain a typo, a personal address one week and a work address the next, or somebody else's address entirely. Every deduplication scheme built on top of it is guesswork.
Update. A log records what was submitted. It has no concept of the current state of a thing. If an application is withdrawn, or an address changes, there is nowhere for that fact to go except a column somebody maintains by hand, which puts a mutable field in an append only structure and creates exactly the overwrite problem the form was avoiding.
Relationships. A database joins tables and enforces that a reference points at something real. Here, an order row referring to a customer contains the customer's name as text, and nothing checks that the customer exists or that the spelling matches. The reconciliation work that follows is the single largest hidden cost of the arrangement.
Constraints beyond the shape of a field. Validation checks that an answer looks like a number or fits a character count. It cannot check that the number is available, that the date is not already taken, or that this person has not already submitted. No question type reads existing data, which rules out every rule that depends on what is already recorded.
None of these are settings that were missed. They are the difference between a log and a database, and no amount of formula work adds them.
The limits that are real, and the ones that are not
Capacity is where the worry usually lands, and it is usually the wrong worry.
A spreadsheet allows up to 20 million cells or 100MB. A twelve column form using a full row per submission reaches that at well over a million submissions, which is far beyond what sends most teams looking for an alternative. Anyone converting an existing file in should note a separate caveat: when a spreadsheet is created from or converted to the Sheets format, any cell with more than 50,000 characters is removed, which matters for long free text answers arriving from elsewhere.
The limits that actually bite are the ones around automation. Scripts, which is how most of the missing behaviour gets patched in, run inside published quotas. Email recipients are capped at 100 per day on a consumer account and 1,500 per day on a Google Workspace account. Total trigger runtime is capped at 90 minutes per day on consumer accounts and 6 hours per day on Workspace. Any single execution is limited to 6 minutes.
Those three numbers decide more architectures than the cell count ever does. A nightly digest and a validation pass fit comfortably. A script that sends an individual acknowledgement on every submission is fine at thirty a day and dead at a hundred and fifty, and it fails without a visible error for the person who submitted. A batch job that rewrites a derived sheet across a large log runs past six minutes and has to be rewritten to process in chunks with its own progress marker, which is real engineering rather than configuration.
Performance is the third practical limit and it is rarely about row count. A log slows down because summary formulas recalculate across the entire sheet on every write. Restructuring so derived figures are computed once into a static range, rather than live across every row, buys back more speed than any amount of archiving.
The mechanical traps worth knowing early
Three failures come up repeatedly and all three are avoidable.
Formula columns and appended rows. Adding a calculation in a column beside the response data works until the next submission arrives, because the new row has no formula in it. Filling the formula down to a thousand rows creates a thousand rows of near empty output that break sorting and counting. The stable pattern is to keep the response range untouched and put all derived work on a separate sheet that reads from it, ideally with a single formula in the header row that expands down the column by itself.
Sorting the response sheet. Sorting or inserting rows inside the range the form writes to is how the mapping between a row and a submission gets scrambled. Sort a copy or a view, never the log.
Header edits. Renaming or reordering columns in the response sheet does not change the form, and the relationship between question and column is positional. A form question added later appends a new column at the end rather than next to its logical neighbours, so a log that has grown over two years has its columns in the order the questions were invented.
There is a fourth, quieter trap. A file upload question stores files in a folder and puts a link in the cell, and to answer one, responders need to sign in to a Google Account. The question type is also unavailable when the form is stored in a shared drive or when an administrator has turned on Data Loss Prevention. A log whose attachments live behind a sign in requirement is not a portable record.
Where to go when it stops fitting
There is no single upgrade path, because the three reasons for outgrowing the arrangement lead to three different places.
| What is actually wrong | What fixes it | What it costs |
|---|---|---|
| Relationships and integrity: orders referring to customers referring to products | A relational database, or a spreadsheet database with linked tables | Modelling work, and somebody who maintains it |
| Volume and query speed beyond a sheet | A database with indexes, reached through a small application | Hosting, backups, an application layer |
| Nothing tracks what happened after a record arrived | A form tool with response management | Less modelling range than a database |
That third row is the one the comparison articles skip, and it is the most common situation. A great many teams reaching for a database do not need joins or indexes at all. What they need is for the submission to carry an owner, a stage, and a record of what was said back to the person who sent it, because those three facts are what currently live in somebody's inbox and head.
This is worth separating clearly. A database answers what is recorded. A tool built around responses answers what is recorded and what has been done about it: the submission arrives with an owner and a stage attached, the reply is composed on the same screen as the answers, every send stays on record with whether it was opened, and contacts assemble themselves from the email address so that one person appearing in four forms remains one person. The pricing model differs in a way that matches the difference: forms and responses are unlimited, and the cost follows the number of people using it. A log that grows every day costs nothing extra; a team that doubles does cost more.
The honest trade is that this category gives up modelling range. If the records genuinely have to reconcile against each other, products against orders against suppliers, a database is the right tool and response management does not substitute for it.
Deciding in one pass
Answer four questions about the data, and the choice usually makes itself.
Does one real world thing appear in several records that must be treated as the same thing? If yes, the missing key is the problem, and either verified email collection or a tool that builds contacts automatically addresses it.
Does any field change after submission? If yes, an append only log is the wrong container for that field, and it belongs somewhere that has a concept of current state.
Does anything need checking against what is already recorded, a remaining place, a date already booked, a second submission from the same person? If yes, the form cannot do it, because no question type reads existing data.
Does the record need to show what happened after it arrived? If yes, and the first three answers were no, then the problem is not that the sheet is not a database. It is that the job has a second half, and that is a different category of tool rather than a bigger table. The features page is the fastest way to see what that second half consists of.
What to change first
Stop putting formulas and hand maintained status columns inside the range the form writes to, and move all derived work to a separate sheet, because that alone removes the most common corruption in this setup. Then answer the four questions above honestly, since three of them point at a database and one points somewhere else entirely. If the answer is the last one, Halict is worth ten minutes against a form you already have.
Q1. How many responses can a linked spreadsheet hold?
A spreadsheet allows up to 20 million cells or 100MB, which works out to well over a million submissions on a typical twelve column form. Slowness usually arrives long before that limit, and it is normally caused by summary formulas recalculating across the whole sheet on every new row rather than by the row count itself.
Q2. Can the same person be recognised across several submissions?
Only as reliably as the field being matched on. If email addresses are collected as verified, they come from the signed in Google Account and are dependable. If they are collected as responder input, the address is whatever was typed, so any deduplication built on it is an estimate. There are no keys and no joins in this arrangement.
Q3. Why do calculations in the spreadsheet break when new responses arrive?
Because the form appends a row that contains only the answers, with nothing in the columns added by hand. Filling formulas down in advance creates blank derived rows that break counts and sorts. Keeping the response range untouched and doing all calculation on a separate sheet that reads from it avoids the problem entirely.
Q4. Can a form refuse a submission based on what is already recorded?
No. Data validation checks the shape of an answer, such as a maximum character count, and no question type reads existing data. That rules out capping a number of places, blocking a date already taken, or rejecting a second submission from the same person. Those checks have to happen after the fact, or in a tool whose form can read its own records.
Q5. Is it worth moving to a real database, or is there something in between?
It depends which property is missing. If records have to reference each other and stay consistent, a database is the right answer. If the real gap is that nothing tracks who owns a submission and what was replied, a database does not close it, because that is correspondence and tracking rather than storage.
