Repository navigation
Generate file checksum on upload #56057
Description
Activity
- added0. Needs triagePending check for reproducibility or if it fits our roadmapPending check for reproducibility or if it fits our roadmap
on Oct 28, 2025 - addedhotspot: file transfer performanceupload & download performance related optimizationsupload & download performance related optimizations
on Oct 28, 2025 @ZetaTom sounds interesting :)
I don't understand why the checksum is needed, couldn't the etag be enough? How is the desktop client doing it?
While relying on checksums could work, for large files, the delay to compute the checksum could add some delay. IIRC, on my machine the md5 checksum take a few seconds for a 1GB file. Depending on the scenario, this may or may not be ok.
I think that asking the server to compute the checksum on every file upload/change would be too much to ask, but maybe we can reduce it to some scenario only?
Reacted by KateRelevant discussions in #11138 and linked issues.
My personal use-case for server-side checksuming would be what's described here; nextcloud/android#255 (comment)
Automatic image upload conflict resolution, when there in reality is no conflict, only an issue of tracking what's already been successfully uploaded. A somewhat recent issue that caused the android app to forget it's sync state (and stop syncing for quite a while) has exasperated this issue for me greatly.From my understanding, currently stored checksums are only available where the client provided one on upload, and the server may not even have verified it.
Computing checksums on upload should be relatively cheap, the file IO tends to be more expensive in my experience, network IO even more so.
Thanks for the added information.
I get the use case, and I would add that this is not really linked to versions, but rather files in general.
As I said, in some scenario, generating the checksum in the server would impact performance. We do have users that are uploading files the size that are 1TB big.
But for your use case, we can probably find a middle ground.
Could something like this work?
If we detect a conflict on a file, the clients request the checksum of the existing file from the server. If the checksum is not known, then the server computes it. The clients can then compare the received checksum with the one of the new file that they can compute locally.
Thanks for the added information.
I get the use case, and I would add that this is not really linked to versions, but rather files in general.
I'm aware that what I'm asking isn't directly related to versions, but I'm providing additional context as to why checksums in general are desirable.
As I said, in some scenario, generating the checksum in the server would impact performance. We do have users that are uploading files the size that are 1TB big.
This may not necessarily be a problem, as you could compute the checksum of the files as they stream over the network, assuming of course that the data arrives in order. (in most checksuming implementations you can feed chunks of data until you are done).
In the case of chunked uploads, you would compute the checksum on assembly (in addition to having verified each chunk beforehand, faster to re-upload a 10MB chunk than the whole 1TB file if something goes wrong).Example of how this can work in Python
In stead of looping over a opened file, streaming chunks from a http client and feeding to the checksumming function at the same time as saving is equally valid.
def checksum(fh, types=['md5', 'sha1']): fh.seek(0) hashes = {} for t in types: hashes[t] = hashlib.new(t) for chunk in iter(lambda: fh.read(4096), b''): for t in types: hashes[t].update(chunk) return dict([(k, v.hexdigest()) for k, v in hashes.items()])
def checksum_file(file_name): chunksize = 16 * 1024 size = 0 md5sum = md5() sha1sum = sha1() sha256sum = sha256() sha512sum = sha512() crc32sum = 0 crc32csum = CRC32CHash() with open(file_name, "rb") as fh: data = fh.read(chunksize) while data != b"": md5sum.update(data) sha1sum.update(data) sha256sum.update(data) sha512sum.update(data) size += len(data) crc32sum = crc32(data, crc32sum) crc32csum.update(data) data = fh.read(chunksize) return { "md5": md5sum.hexdigest(), "sha1": sha1sum.hexdigest(), "sha256": sha256sum.hexdigest(), "sha512": sha512sum.hexdigest(), "crc32": f"{crc32sum & 0xffffffff :08x}", "crc32c": crc32csum.hexdigest(), "size": size, }
But for your use case, we can probably find a middle ground.
Could something like this work?
If we detect a conflict on a file, the clients request the checksum of the existing file from the server. If the checksum is not known, then the server computes it. The clients can then compare the received checksum with the one of the new file that they can compute locally.
Yes, that'd be ideal for my use case, and likely solve a few android sync issues.
I would argue that checksum verification is doubly important on such large files, but I agree that computing the checksum on demand would indeed be a valid option.
Also having the option to verify file integrity of a certain path based on stored checksums would be good (occ command or similar).
(this is all slightly off topic for this particular enhancement suggestion, sorry for hijacking the issue)
Reacted by Narcis GarciaIn the case of chunked uploads, you would compute the checksum on assembly (in addition to having verified each chunk beforehand, faster to re-upload a 10MB chunk than the whole 1TB file if something goes wrong).
That's a good point, but nowadays for S3 we do the final assembly on the S3 itself.
In any case, for most files, we are not using multipart upload, and computing a checksum when writing to disk could be ok. If it is, the clients would be able to have a checksum for most files in the
PROPFINDrequests, and would need to request it very rarely.Any opinion on which checksum to use? I would naively pick the fastest as there are no security implications, and I don't expect collisions at the scale of a single file to be probable.
sha256 is the current defacto standard for the task afaik, but sha1 should be fine. md5 is still used a fair bit, but if there is only one hashing algorithm, I'd want at least sha1.
In the issue(s) I linked, people wanted both md5 and sha1 at the same time, for better compatability with other services who also provide checksums, for syncing purposes.
Even if picking only one algorithm right now, I would strongly recommend making provisions for future support for configurable multiple algorithms. I don't think that would significanttly complicate the code, but would make future work in adding for example md5 and sha256 in addition to sha1 in the future, should the need arise.
Personally I would be more than happy to trade a few CPU cycles for data integrity and compatability in this regard, others may want just some simple integrity checks.
I think being able to pre-generate file hashes the same way as thumbnail from the cli would be great, same with verifying (scrubbing) disk data against stored hashes.
Reacted by Louis and imShaikhAR- changed the title
[-]Include Checksums in File Versioning[/-][+]Generate file checksum on upload[/+]on Nov 5, 2025 xxHashwas also suggested when running the idea through other engineers. Mostly for performance reasons.This also was linked: https://php.watch/articles/php-hash-benchmark
I honestly do not really get why a checksum here is needed in the first place.
For integrity between client and server, the transport protocol guarantees it.
For conflicts in files the etag is exactly doing it, it should change when the file was changed?I'm not familiar with xxHash, nor how widely available it is in different programming languages - I'd strongly recommend going with something that would be easy to implement in a variety of programming languages, preferably with standard library functionality.
How are etags made? Are they checksums, or essentially random UUIDs?
The problem I am facing would not be solvable with a UUID, as I started with a copy from a different system before enabling sync from my phone. This was further exasporated by a client bug a while back which seems to have made the android client forget it's sync settings and state.
Unless it is based on a hash of the content and could be deterministically reproduced on a different system with a different programming language, it is not suitable for resolving sync conflicts.
I have also seen mentions of files getting corrupted with sections of null bytes, this could only be uncovered by checksums.
While TCP packets are covered by a 16bit checksum, that doesn't guarantee there will never be any issues with file transfers in any way, and covers none of what happens at the application layer.
Neither a randomly generated file ID, nor transport layer crc would uncover a partially transferred file, a file checksum could.
I currently have almost 1500 image upload conflicts in the android app, getting repeatedly retried and spamming conflict notifications and thus unnecessarily burning phone battery. Manually verifying and resolving all these files is not feasible.
A large part of them come from the previously mentioned bug messing up sync, but previously I also saw this happening regularly, especially when my phone switched networks soon after taking a picture (walking out of or into WiFi range would somehow interrupt the upload process in such a way that a file was created on the server, but the client was not made aware of it). I don't know if those transfers where fully complete, and just the callback to mark as transfered was interrupted, or if those files now are truncated server side.
Server checksums that the client can use to check local files would allow clients to verify whether the file in full is on the server, without temporarily downloading said file and comparing byte for byte.
A client checksum getting verified by the server before storing the file at the requested path would allow discarding partial/interrupted/unsuccessful transfers instead of assuming the file is good.
Reacted by Louis and rmelotteThank you all for the constructive comments and suggestions!
As for me, the paper I'm working on is about improving and implementing robust bidirectional synchronisation for the Nextcloud Android client.
The problem with Etags is that they are by definition opaque and any given Etag might not relate to the actual file content but may instead be just a random string.
As far as I can tell Etags in Nextcloud are generated by the server by concatenating various properties of the file and hashing them using md5. If I understand correctly,
mtime,ino(inode number?),dev(device id?), andsizeare properties which are hashed to calculate the Etag. See/lib/private/Files/Storage/Local.phpand please correct me if I'm wrong.This causes a few issues. For one clients have no way of locally generating or predicting an Etag of a given file. The only way for a client to generate an Etag is to upload the file and query the server for the resulting Etag. One use case where this doesn't work is (unidirectional) auto upload.
Clients may have files locally that are already on the server. This may happen when a user chooses to reinstall the app or after the user manually synced their photos via USB, for example. In these cases there is no way for the client to check if the file on the server is in fact the same as the local one, so clients have to download the whole file and check if it's the same or blindly replace the one already on the server.
For bidirectional sync, where both ends can perform changes on the files, this issue is becomes even more apparent. Currently, the Android client stores the last known Etag of the file on the server. When a local change is detected, this record is cleared. During the next sync, we compare this record to the one returned by the server. If the Etag on the server has changed but the 'local Etag' hasn't, the file from the server will replace the local one. On the other hand, if the 'local Etag' has been changed and the one on the server hasn't, the local file is uploaded to replace the one on the server. This works great until both Etags change. In that case we have no other option, other than to show the dreaded conflict resolution dialogue. In the past there have been several instances where the local database had to be reset along with all the 'local Etags'. This causes the app to identify all files as 'in conflict'. Users having a few thousand photos in their Nextcloud will probably not bother checking every single file. This may also happen when users choose to reinstall the app or migrate to another device.
In these cases Etags are not enough given their opaque nature. It is entirely possible to have different Etags for the same file, where the content, and therefore the checksums, are identical. We also observed a few instances where the Etags on the server have been reset, either causing all files to be redownloaded by the clients, or to be marked as conflicts. Per definition Etags should change as soon as the data returned by a GET query changes. There are also other problems with relying on Etags for synchronisation; as far as I can tell there is provision for the same Etag to be returned, when a change was undone, for example. These are all things that have happened and frustrated me as a user in the past.
To address these use cases a simple checksum for individual files would suffice to prevent most rouge conflicts from appearing in the first place.
How this will be implemented and when the calculation should be performed needs to be carefully considered. When performing the calculation after a change has been detected or an upload was performed, the checksum is immediately available and can be queried from clients without too much delay. This may, however, cause severe load spikes on the server, when a user returns from a trip where they recorded a lot of videos, for example. On the other hand, if calculation is performed on-demand, clients may have to wait a substantial amount of time for each file that needs to be synchronised.
Regarding the time spend calculating the checksum: in my opinion the benefits of having a concrete metric which guarantees both integrity and more efficient conflict resolution outweighs the downsides. For users uploading particularly large files, I would agree with @oddstr13 comments, that checksums are a valuable asset, especially when performing such large operations. The overhead of manually checking that the file is indeed what it's supposed to be, or having the client itself re-download the whole file just for a sanity check, I think is much greater than calculating a checksum.
During my research I performed a few basic benchmarks. As a baseline I used the algorithms listed as supported by the
<oc:checksums />property. All of these tests were run on an AMD Ryzen 5 7640U. The data which is begin hashed was the output ofdd if=/dev/urandom of=random_data.bin bs=1M count=1000. All tests were run 100 times and their averaged results can be found in the following table.Algorithm Real Time MD5 1.101 s SHA1 1.049 s SHA256 2.249 s SHA3-256 1.781 s Adler32 0.006 s I know that these results are not representative but it illustrates the point that choosing a suitable algorithm for this use case is important; to minimise overhead and provide enough resolution, as to also minimise the probability of hash collisions, especially for smaller files. We need to keep in mind that the same algorithm needs to perform well on the server and on low powered mobile devices, where resources are especially limited.
While the introduction of checksums would allow us to eliminate most of the rouge conflicts, some still remain. Addressing these and minimising the need for manual conflict resolution is the focus of the paper I'm working on.
The main strategy to further reduce apparent conflicts is to introduce a new relation between files and their ancestor versions. Consider that the current version of a file (f) is the result of many previous versions of that file (A(f)) which were slightly different in some ways. If we now have a history of these files, or changes, we can use that history to assist in conflict resolution.
We're trying to do essentially that in the Android client, by remembering the last Etag and comparing that to the current one. What I'm trying to do is to take that idea further by extending the history which is used to infer information about which file to keep in case of a conflict between different versions of the same file. If we extend this history, we also increase the amount of information available to automatically resolve conflicts in the background without requiring any user interaction.
Consider the scenario outlined above, where a file (f) that has been synchronised across multiple device underwent a change (to become f') which has later been made undone (f''). In this case we could examine the history and see that A(f'') incudes all the versions of A(f) and a few more edits on top. In this case we could assume that f is outdated and should be replaced by f''. This is a very niche application but it does allow us to take care of conflicts which would be difficult to resolve automatically without this additional bit of information.
Calculating the checksums for file versions could be as easy as persisting the current checksum along with the file version. If we opt for on-demand calculation, this becomes a bit more difficult but could also be done asynchronously, for example, when a new file version is being created on the server.
In my paper I've outlined a set of conflict resolution policies which would greatly improve automatic conflict resolution for the Android client. I have to say that I'm approaching my due date and the end goal is to investigate ways to do efficient bidirectional sync on mobile clients and it will probably not result in a production ready solution, but rather a proof of concept.
@artonge, this issue is specifically about adding the option to support checksums of individual file versions, as outlined in the description. The current title of the issue, unfortunately, doesn't reflect that fact.
Before opening this issue, I've found the following one opened by @tobiasKaminsky, which was already mentioned by @oddstr13 in a previous comment, that specifically requests the availability of checksums to verify up- and downloads: #11138
Additionally there's already an issue open requesting the option to verify chunked uploads via a checksum mechanism, also opened by @tobiasKaminsky: #13509
Reacted by imShaikhAR and Narcis Garcia@artonge, this issue is specifically about adding the option to support checksums of individual file versions, as outlined in the description. The current title of the issue, unfortunately, doesn't reflect that fact.
In my opinion, this should still be tied with general file management, and not only files_versions.
And for your information, this was discussed yesterday in the context of auditing files operations in an internal call. This has good chances of being implemented not so far in the future.
Reacted by imShaikhAR and Denis DordoigneReacted by rmelotte, Narcis Garcia and Denis DordoigneWhen using S3 as a storage - maybe the checksum computation can also be offloaded to S3.
Reacted by Louis
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsTriaged
How to use GitHub
Is your feature request related to a problem? Please describe.
While working on nextcloud/files-clients#42 and nextcloud/android#19, I came up with an experimental concept to improve synchronisation with mobile clients and introduce robust bidirectional synchronisation for the Android client.
For this concept to work, the server needs to provide a history of the files in question. This can be provided via the
files_versionsapp. Another requirement is the introduction of server-side checksums. This allows clients to locally check the current file against a backlog of versions to improve the handling of file conflicts and minimise manual intervention.Describe the solution you'd like
To be concrete: I would like the versioning record for each file version to be extended to include a checksum of the actual file content. The checksum can either be calculated when a file version is created (persisted to memory), or on demand, when a client queries the corresponding property.
Ideally the current response would be extended to the following:
Including
<oc:checksum>for any individual version.Describe alternatives you've considered
Several alternative approaches have been considered. These would entail major changes to the way mobile clients interact with the server and have thus been ruled out.
Additional context
Server-side checksums can already be requested by querying the
<oc:checksums />property of files. Clients can also request that the server hashes a given file on demand usingChecksumUpdatePlugin.The implementation of this concept is part of an academic project. Any comments, discussions or contributions on the implementation or viability of the proposed solution are welcome.