Repository navigation
Files search takes a long time with unified search #23835
Description
Activity
- added1. to developAccepted and waiting to be taken care ofAccepted and waiting to be taken care of
on Nov 2, 2020 Also, there is no x in the input field during loading so it’s impossible to cancel the search – you have to wait for it to finish to start a different one. To fix this, we could replace the spinner with an x on hover/focus.
Searching cancel other searches.
You can always change your query and it should trigger the search againIt's mainly slow because the files search implementation doesn't paginate. So you'd always build the full results list on the back-end before it's chopped into the individual pages. This doesn't scale, as you observed. Our instance just has a lot of files and shares, hence the slowness.
Reacted by John Molakvoæ, YANO Tetsuro, Emilian Mitocariu, Chris Puttick and neveroils@ChristophWurst non-pagination is NOT the problem:
Pasting
152151into the search term will yield a single result after ~4.7s (runtime taken from Firefox devtools).Executing a query directly on the PostgreSQL 12.0 server (2.6e6 rows in oc_filecache):
SELECT * FROM oc_filecache WHERE name LIKE '%152151%: 0.5sSELECT * FROM oc_filecache WHERE name ILIKE '%152151%: 1.1s
This makes NC file search at least 4 times slower than the database full table search.
If the search term is not entered by copy/paste, but typed in, things get much worse, since each character will trigger a new query on the server, consuming more resources (when typing fast, the final search request will take 16s).
@skjnldsv A new search does not cancel the running queries in the database backend.
Reacted by Christoph Wurst, YANO Tetsuro, Chris Puttick, Adam Monsen, neveroils and brtptrs@skjnldsv A new search does not cancel the running queries in the database backend.
Right, it cancel the frontend request of course.
I enabled log_min_duration_statement=800 in the PostgreSQL Server, and found that each search will execute the same query four times, which explains the x4 observed above.
@ChristophWurst non-pagination is NOT the problem:
If there was pagination the query would include more
WHEREclauses to scope the range. It would therefore help.But yeah,
ILIKEis expensive.Not really, if the result set is small, or not at all when using window functions, since the time consuming part is the sequential scan, not the result set transfer.
For PostgreSQL, speed could be improved by not using
name ILIKE :searchExpression(1.1s), butlower(name) ILIKE lower(:searchExpression)(0.8s) or even betterlowercasedname LIKE lower(:searchExpression)(0.5s).But there are fruits hanging much lower:
- eliminating those 3 exceeding queries, there's something going awfully wrong here.
- delaying search for some 300-500ms after the last keystroke, to prevent multiple zombie queries eating database performance.
More advanced:
- split filesearch in two requests: First search the current directory only (my example has 4800 Files in the folder, <20ms), then the full scan. Quite often, only the current directory is of interest.
Reacted by neveroils and benedettoI found the reason why the database was hammered with the same query 4x for a single search. They correspond with four shares from the same user. These result in four MountPoints, having different different roots on the same physical storage (same numericId). Unfortunately, the filtering code is buried very deep in CacheWrapper, and each wrapper retrieves the full data from its Cache instance before filtering.
Hiding (parts of) the filtering from the View in separate Caches might be good programming practice in general, but is a real pain for big datasets. Can't see a way how to speed this up in a non-dirty way.Suggest re-titling this issue in line with #24029 , since this is a regression in behavior from NC19. At best, we have lost the "fast search" capability. At worst, we've lost "search" alltogether.
My personal use case: updated to NC20 last night. Today, any search query will peg the mysql server and make nextcloud unresponsive. Having just performed a
files:cleanmyoc_filecachetable has 2057547 rows. It takes 2 seconds just to run the count query. :|edit: it works... just very very slowly. > 10 seconds to search
testing. So far 280s forlarge files for youtube. In mariadb processlist I see:MariaDB [nextcloud]> show full processlist; +-----+-------------+---------------------------------+-----------+---------+------+--------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+----------+ | Id | User | Host | db | Command | Time | State | Info | Progress | +-----+-------------+---------------------------------+-----------+---------+------+--------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+----------+ | 2 | system user | | NULL | Daemon | NULL | InnoDB purge worker | NULL | 0.000 | | 1 | system user | | NULL | Daemon | NULL | InnoDB purge coordinator | NULL | 0.000 | | 3 | system user | | NULL | Daemon | NULL | InnoDB purge worker | NULL | 0.000 | | 4 | system user | | NULL | Daemon | NULL | InnoDB purge worker | NULL | 0.000 | | 5 | system user | | NULL | Daemon | NULL | InnoDB shutdown handler | NULL | 0.000 | | 724 | nextcloud | nextcloud.compose_default:52848 | nextcloud | Query | 391 | Sending data | SELECT `filecache`.`fileid`, `storage`, `path`, `path_hash`, `filecache`.`parent`, `name`, `mimetype`, `mimepart`, `size`, `mtime`, `storage_mtime`, `encrypted`, `etag`, `permissions`, `checksum`, `metadata_etag`, `creation_time`, `upload_time` FROM `oc_filecache` `filecache` LEFT JOIN `oc_filecache_extended` `fe` ON `filecache`.`fileid` = `fe`.`fileid` WHERE (`storage` = 9) AND (`name` COLLATE utf8mb4_general_ci LIKE '%test%') | 0.000 | | 734 | nextcloud | nextcloud.compose_default:52950 | nextcloud | Query | 391 | Sending data | SELECT `filecache`.`fileid`, `storage`, `path`, `path_hash`, `filecache`.`parent`, `name`, `mimetype`, `mimepart`, `size`, `mtime`, `storage_mtime`, `encrypted`, `etag`, `permissions`, `checksum`, `metadata_etag`, `creation_time`, `upload_time` FROM `oc_filecache` `filecache` LEFT JOIN `oc_filecache_extended` `fe` ON `filecache`.`fileid` = `fe`.`fileid` WHERE (`storage` = 9) AND (`name` COLLATE utf8mb4_general_ci LIKE '%testing%') | 0.000 | | 746 | nextcloud | nextcloud.compose_default:53052 | nextcloud | Query | 390 | Sending data | SELECT `filecache`.`fileid`, `storage`, `path`, `path_hash`, `filecache`.`parent`, `name`, `mimetype`, `mimepart`, `size`, `mtime`, `storage_mtime`, `encrypted`, `etag`, `permissions`, `checksum`, `metadata_etag`, `creation_time`, `upload_time` FROM `oc_filecache` `filecache` LEFT JOIN `oc_filecache_extended` `fe` ON `filecache`.`fileid` = `fe`.`fileid` WHERE (`storage` = 9) AND (`name` COLLATE utf8mb4_general_ci LIKE '%testing%') | 0.000 | | 760 | nextcloud | nextcloud.compose_default:53164 | nextcloud | Query | 320 | Sending data | SELECT `filecache`.`fileid`, `storage`, `path`, `path_hash`, `filecache`.`parent`, `name`, `mimetype`, `mimepart`, `size`, `mtime`, `storage_mtime`, `encrypted`, `etag`, `permissions`, `checksum`, `metadata_etag`, `creation_time`, `upload_time` FROM `oc_filecache` `filecache` LEFT JOIN `oc_filecache_extended` `fe` ON `filecache`.`fileid` = `fe`.`fileid` WHERE (`storage` = 9) AND (`name` COLLATE utf8mb4_general_ci LIKE '%arge fil%') | 0.000 | | 769 | nextcloud | nextcloud.compose_default:53284 | nextcloud | Query | 320 | Sending data | SELECT `filecache`.`fileid`, `storage`, `path`, `path_hash`, `filecache`.`parent`, `name`, `mimetype`, `mimepart`, `size`, `mtime`, `storage_mtime`, `encrypted`, `etag`, `permissions`, `checksum`, `metadata_etag`, `creation_time`, `upload_time` FROM `oc_filecache` `filecache` LEFT JOIN `oc_filecache_extended` `fe` ON `filecache`.`fileid` = `fe`.`fileid` WHERE (`storage` = 9) AND (`name` COLLATE utf8mb4_general_ci LIKE '%arge file%') | 0.000 | | 771 | root | localhost | nextcloud | Query | 0 | Init | show full processlist | 0.000 | +-----+-------------+---------------------------------+-----------+---------+------+--------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+----------+ 11 rows in set (0.001 sec)Reacted by Chris Puttick and Lorenzo P.#25136 should help with search performance, how much of a difference it makes will depend greatly on setup specifics.
135 remaining items
Hi, please update to at least 23.0.12 and report back if it fixes the issue. Thank you!
- added0. Needs triagePending check for reproducibility or if it fits our roadmapPending check for reproducibility or if it fits our roadmapand removed1. to developAccepted and waiting to be taken care ofAccepted and waiting to be taken care of
on Nov 26, 2022 I think this is good on 25 already.
Thanks for testing!
For what it's worth, I think it was also good on 24, as I was running it for a few weeks before 25 was stable.
Thanks for checking.
For those that have thousands of files and potentially many search results with each search, displaying five at a time is not an ideal solution. Does such merely mask the issue?
I agree with the comment about the UI and only showing five results at a time, but it is unrelated to the performance issues or fixes. The performance issues were there when the UI was already only showing 5 results at a time, unless I'm mistaken.
Nothing changed there. It's just the performance that improved.
Edit: still, I'd be happy with a way to open search results in a proper, long list format. Like it used to be before it got reduced to showing results in the search panel alone.
Edit: still, I'd be happy with a way to open search results in a proper, long list format. Like it used to be before it got reduced to showing results in the search panel alone.
I second this. The current small search panel is not very user friendly. It's small and doesn't show the full names/file paths of files when you search, and is clunky to scroll through. Furthermore, if you type a search then hit enter ↩, it'll automatically open the first result, which is counter-intuitive.
It'd be great to have a full search page available. Gmail gives a good example of good search usability. Showing a popup window initially, and then another window if you hit enter.
Reacted by pjft and Sashko TodorovCan I second this too ? second seconder? or a thirder?
The UI has some room for improvement.
Depending on the amount of files, the file search component of unified search takes a long time to get results back. It’s the slowest of all – Talk, Mail and everything loads before. On our own instance, it takes roughly 40 seconds.
Also, there is no x in the input field during loading so it’s impossible to cancel the search – you have to wait for it to finish to start a different one. To fix this, we could replace the spinner with an x on hover/focus.
Gif for illustration, showing how it takes 40 seconds:

(This is very related to the direct filtering of the current folder, which also doesn’t work: #23432 )
cc @nextcloud/search