The text is not the paper
A few notes on authorship in Big Life Sciences
None of the below are strict rules, they are cultural norms and practices.
This was motivated by last week’s discussion on authorship and responsibilities. I think that there was a lot of talking past each other because for some people it was not obvious how you could be a co-author on something that you had not, in any meaningful sense, written.
Thus, I thought it would be useful to write down some observations about publications that involve large collaborations, which I am going to call Big Life Sciences. These are partly answering some of the questions that came up (mostly on Twitter) and partly more general observations.
The text is not the paper
In the sciences, and certainly in Big Life Sciences, the text is not the paper.
One can be a co-author without having written a single word of text because the paper is the data, the analyses, the ideas, the figures, and (also) the text. The paper may even include elements such as code on Github, a website, data on a public repository. These are all important parts of the paper and the text may not even be the most relevant element. Thus, one can legitimately be a co-author without having written any of the text (by having contributed to the data analysis or the figures or data collection).
Everyone is a co-author, but some are more co-author than others
The author list can be broken down into three categories: the first author(s), the middle authors, and the last author(s). The first and last authors are the most important and get the most credit, while the middle authors did some crucial work but are not the main drivers of the project.
Wherever you are reading this, there is probably a McDonald’s within walking distance (or at worst, a short drive away). That McDonald’s is, however, not owned by the McDonald’s corporation, but rather by a local businessperson. This businessperson makes a lot of money. Why does the McDonald’s corporation allow this? Why is your local McDonald’s not managed by a mid-level manager who is an employee of the McDonald’s corporation? Surely they could capture these profits.
Fast food operates on low wages and low margins. When the teenager who was supposed to work on Sunday morning calls in sick and the shift is short-staffed, the owner will step in to operate the cash register and mop the floor. They will do so because it’s their McDonald’s. A middle manager at a large corporation, on the other hand, is not going to step in, leave their family in the lurch on a weekend, to go in to work and mop the floors.
Typically, the first author is the junior person (student or postdoc) who did most of the actual work and the last author is their supervisor. The middle authors are other people who contributed to the project, but not as much as the first author. The first author is not just the person who did the most work, but also the person whose McDonald’s it is. When the journal administratively rejects the paper and you need to resubmit because the figure files did not adhere to the right naming convention (they need to be called Smith_Figure1.tif and not Figure1.tiff), the first author is the one who will be doing the work to fix the problem and resubmit. If a co-author is flaky, the first author is the one who will be doing the work to chase them down and get them to do their part. It is their McDonald’s.
The last author is the corresponding author (technically these are not the same, but let’s keep it simple). They take responsibility for the paper and act as the point of contact for others to ask questions.
You can have multiple first and multiple last authors, but some are firster and laster. This is signified by an asterisk next to the names of the first authors, and a footnote that says “These authors contributed equally to this work”. Similarly, multiple last (or corresponding) authors can be listed. This is a way to give credit to multiple people who contributed significantly to the work. In practice, the first first author usually gets more credit than the second (or third) first author, and the last last author usually gets more credit than the second (or third) last author. This often emerges from collaborations between different labs where there is not clear hierarchy of who is the first and last author, but can also emerge from collaborations within the same lab where there are multiple people who contributed significantly to the work.
Middle authors get less credit than the first and last authors, but they can be important contributors to the project. In fact, collectively, the middle authors will often have contributed most of the paper!1 They may have contributed specific analyses or specific expertise that the main authors are unable to perform themselves. This is the main reason for collaboration and interdisciplinary work. Bioinformaticians have long bemoaned how they are often crucial contributors to a project but get relegated to middle authorship which is not adequately rewarded. Middle-author disease is a term I have heard from many.
How does such a paper get written in practice?
There is a lot of variability in this process and unlike the question of what it means to be first author, there less of a cultural consensus. I have seen a Google Doc that is shared from the start that everyone can see and edit. At the other extreme, I have seen a first author ask individual co-authors for bits and pieces, then stitch them together before sharing what is very close to a final draft with everyone.
In my group, I tend to advocate for the following: first author writes first draft, which is shared with me and I start to co-edit. Then, we progressively share the draft with individual co-authors, folding in their comments before sharing with the next co-author. You only get someone to read it for the first time once, and those cold reads are very valuable. Eventually, everyone will have seen it and provided comments, and we’ll have one final call for comments before submission.
In any case, it is very easy for someone to end up in a situation whereby they have legitimately contributed to a paper and should, by the norms and ethical standards of the field, be a named co-author, but they have not seen the text until the very end and may even not fully understand all of it. One of the advantages of collaboration is that you can have people with different areas of expertise contributing to a joint effort. This is not per se a problem, but it does cut against strict liability for all co-authors.
Should different sections of the paper be written and signed by different authors?
This was one of the suggestions that cropped up in discussion. After a bit of thinking, I actually believe that this is not a good idea. In as much as there is value in having a coherent structure and voice across a paper, it is harmful to have too many cooks in the kitchen. For example, in our manuscripts, we have decided to use the term habitat to describe different places microbes live. This was the result of a lot of discussion and debate, and we have decided that it is the best term to use. If we had different people writing different sections of the paper, we might end up with some sections using the term habitat and others using the term environment or niche. This would make the paper less coherent and more difficult to read.
When we have everyone edit together, it is easy to end up with duplications or unfired Chekhov’s guns (terms that are introduced and then never mentioned again). Also, people are generally better at identifying problems than solving them, so often co-authors will provide an edit that correctly pinpoints a problem, but their solution is not the best. Thus, it is my experience that the best results are when you combine a large number of people providing input and feedback, but a core group of people who make the final decisions to maintain coherence.
I acknowledge that there is a tension between accountability and coherence, but I think it is best to accept a small loss of accountability in order to get better papers overall.
If you have, say 22 authors; that’s 1 first and 1 last plus 20 in the middle. Each individual in the middle may have done relatively little, but their contributions may be, in aggregate, larger than the first author.


