A deadlock that has been found on our Community servers is:
```
------------------------
LATEST DETECTED DEADLOCK
------------------------
2022-03-03 06:14:37 0x2aec58887700
*** (1) TRANSACTION:
TRANSACTION 1043681684, ACTIVE 0 sec starting index read
mysql tables in use 1, locked 1
LOCK WAIT 23 lock struct(s), heap size 1136, 137 row lock(s), undo log entries 26
MySQL thread id 448011, OS thread handle 47195956565760, query id 752406931 10.128.147.211 mmcloud updating
DELETE FROM SidebarChannels WHERE ((ChannelId IN (?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?) OR ChannelId IN (?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?)) AND CategoryId = ?)
*** (1) WAITING FOR THIS LOCK TO BE GRANTED:
RECORD LOCKS space id 907 page no 368 n bits 0 index PRIMARY of table `mattermost`.`SidebarChannels` trx id 1043681684 lock_mode X waiting
Record lock, heap no 121 PHYSICAL RECORD: n_fields 6; compact format; info bits 0
0: len=26; bufptr=0x2aebbd36b75c; hex= 3363776f6a6a6d36316a62663967657036706872316533696a77; asc 3cwojjm61jbf9gep6phr1e3ijw;;
1: len=26; bufptr=0x2aebbd36b776; hex= 3568616279726d36746a6265747061656538786973786d6b7177; asc 5habyrm6tjbetpaee8xisxmkqw;;
2: len=26; bufptr=0x2aebbd36b790; hex= 373568726a7a68676169726370653477696877727866716f7463; asc 75hrjzhgaircpe4wihwrxfqotc;;
3: len=6; bufptr=0x2aebbd36b7aa; hex= 00003e355176; asc >5Qv;;
4: len=7; bufptr=0x2aebbd36b7b0; hex= 39000352420fce; asc 9 RB ;;
5: len=8; bufptr=0x2aebbd36b7b7; hex= 80000000000000fa; asc ;;
*** (2) TRANSACTION:
TRANSACTION 1043681679, ACTIVE 0 sec starting index read
mysql tables in use 1, locked 1
LOCK WAIT 76 lock struct(s), heap size 1136, 294 row lock(s), undo log entries 72
MySQL thread id 447990, OS thread handle 47194812454656, query id 752406930 10.128.147.211 mmcloud updating
DELETE FROM SidebarChannels WHERE ((ChannelId IN (?,?,?,?,?,?,?,?,?,?,?,?,?,?,?) OR ChannelId IN (?,?,?,?,?,?,?,?,?,?,?,?,?,?,?)) AND CategoryId = ?)
*** (2) HOLDS THE LOCK(S):
RECORD LOCKS space id 907 page no 368 n bits 0 index PRIMARY of table `mattermost`.`SidebarChannels` trx id 1043681679 lock_mode X
Record lock, heap no 121 PHYSICAL RECORD: n_fields 6; compact format; info bits 0
0: len=26; bufptr=0x2aebbd36b75c; hex= 3363776f6a6a6d36316a62663967657036706872316533696a77; asc 3cwojjm61jbf9gep6phr1e3ijw;;
1: len=26; bufptr=0x2aebbd36b776; hex= 3568616279726d36746a6265747061656538786973786d6b7177; asc 5habyrm6tjbetpaee8xisxmkqw;;
2: len=26; bufptr=0x2aebbd36b790; hex= 373568726a7a68676169726370653477696877727866716f7463; asc 75hrjzhgaircpe4wihwrxfqotc;;
3: len=6; bufptr=0x2aebbd36b7aa; hex= 00003e355176; asc >5Qv;;
4: len=7; bufptr=0x2aebbd36b7b0; hex= 39000352420fce; asc 9 RB ;;
5: len=8; bufptr=0x2aebbd36b7b7; hex= 80000000000000fa; asc ;;
[bitmap of 256 bytes in hex: 00 00 00 00 00 00 00 f0 00 00 00 00 00 00 00 02 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 ]
*** (2) WAITING FOR THIS LOCK TO BE GRANTED:
RECORD LOCKS space id 907 page no 67 n bits 0 index PRIMARY of table `mattermost`.`SidebarChannels` trx id 1043681679 lock_mode X waiting
Record lock, heap no 103 PHYSICAL RECORD: n_fields 6; compact format; info bits 0
0: len=26; bufptr=0x2aeb4515ec1b; hex= 6431376b35706f71786979776d6b6d616b6a78623539356b366f; asc d17k5poqxiywmkmakjxb595k6o;;
1: len=26; bufptr=0x2aeb4515ec35; hex= 7067653939686a356562676b38636f6f70637a6e7871677a6568; asc pge99hj5ebgk8coopcznxqgzeh;;
2: len=26; bufptr=0x2aeb4515ec4f; hex= 336e3731636a73656f667931397267746664337962786d713777; asc 3n71cjseofy19rgtfd3ybxmq7w;;
3: len=6; bufptr=0x2aeb4515ec69; hex= 00003e34fa4d; asc >4 M;;
4: len=7; bufptr=0x2aeb4515ec6f; hex= 4600026d940e4c; asc F m L;;
5: len=8; bufptr=0x2aeb4515ec76; hex= 8000000000000082; asc ;;
*** WE ROLL BACK TRANSACTION (1)
```
It seems like 2 DELETE queries are somehow deadlocking on the same row.
On looking at it deeply, we can see that in an updateSidebarCategory scenario,
we run 2 delete queries in a single transaction. In that delete query, we not only delete the existing channels for a category,
but also delete the new channels that are going to be inserted in the new category.
The potential issue with this is that since we are updating both the source and destination category in a single transaction,
the category which is moved appears in the destination section for the first category, and in the source section for the second category.
Which means in a single transaction, you have 2 queries possibly interlocking due to the same channels.
Therefore, if we run 2 separate updateCategories with an inverted order of categories, then the same set of delete queries
can lock on the opposite order of channels and cause a deadlock.
I haven't been able to reproduce this, but just a hunch. And in any case, this will reduce the number of DB updates.
```release-note
NONE
```
Co-authored-by: Mattermod <mattermod@users.noreply.github.com>
* Added Archived column to FileInfo
* fixed typo
* Added new column to store column list
* Fixing tests
* Updated migration number to syncup with master
Just a quick POC to move fast :P
We use a search pointer to keep track of
the next row to inesrt to. For every new search
we increment the pointer and do modulo 5.
This means that the value will always remain
between 0-4. And that way, we will always overwrite
the oldest entry on every search.
And while getting the results, we get
all results for that user.
The search parameters are json marshalled
and stored as a JSON blob. This is because
there is no need to search/filter them
in the DB.
Pending items:
Tests obviously.
To improve:
The client needs to send the channel ids
instead of channel names.
We add 2 new params to channel members query.
1. Filter by teamId.
2. Negate that filter.
We include some more optimizations like:
- Moved the team role checks inside the dataloader.
- Moved the channel pretty name computation inside the loader.
Now that we load less data on initial load, we can reduce
the concurrency requirement to be a bit on the safer side.
```release-note
NONE
```
Because of the fact that t.Unix() family of methods do not
contain the monotonic time, there is no guarantee that
consecutive methods will increase in time.
This is more of a best effort to double the time slept,
but in reality there is no way to control this unless
you specifically control your servers, which is hard to achieve
in a CI environment.
https://mattermost.atlassian.net/browse/MM-43848
```release-note
NONE
```
Old versions of the Mattermost server did not qualify queries scanning both `Posts` and `Threads`, and choke on the ambiguity in deciding between the new `DeleteAt` on `Threads` and the `DeleteAt` on `Posts` in existing queries.
While this problem is transient only while running multiple server versions, it effectively makes our backwards compatibility guarantee void, not to mention complicating cloud deployments.
Work around this by renaming `Threads.DeleteAt` to `Threads.ThreadDeleteAt`. The old migration is nulled out, but remains, since some test servers have already upgraded and manually fixing each affected instance would be problematic. Thew new migration takes care of removing the old column -- if it ever existed.
Fixes: https://mattermost.atlassian.net/browse/MM-43770
Co-authored-by: Mattermod <mattermod@users.noreply.github.com>
* [MM-42739] Initial setup for top channels for team
* [MM-42739] Add initial tests
* [MM-42739] Update tests
* [MM-42739] Add top channels for user
* [MM-42739] Fix query
* [MM-42739] Update query
* [MM-42739] Improve query performance
* [MM-42739] Remove rank
* [MM-42739] Fix tests to use new time range today
* [MM-42739] Add tests for top channels for user
* [MM-42739] Add test for pagination
* Remove top channels by time struct
* [MM-42739] Update test names
* [MM-42739] Remove rank from top reactions
* [MM-42739] Return empty array instead of nil when result is empty
* [MM-42739] Add additional tests and update permissions check for teams
* [MM-42739] Add excluded channel tests for top reactions
* [MM-42739] Move insights to api4/insights and keep time range as string until required
* [MM-42739] Update queries only check DeleteAt after union
* [MM-42739] Improve query performance by using publicchannels table
* [MM-42739] Fix broken query after merge
* removed appending the root post to the posts list
also changed `UpdateAt` to `CreateAt` in thread_store.go in accordance with kyriakos
* fixed failing test after latest change
Co-authored-by: Mattermod <mattermod@users.noreply.github.com>
* Disambiguates some units.
* Updates DB attribute.
* Updates some more error-prone units.
* Updates some tests with legible constants.
* Updates query for MySQL case sensitivity.
* Fixes more casing issues.
Co-authored-by: Mattermod <mattermod@users.noreply.github.com>
On Postgres, `GetTeamsUnreadForUser` triggers a sequential scan on `Posts`. We can avoid this by querying the `Threads` table directly and only joining to `Posts` to eliminate deleted threads. (We could avoid the latter if we later denormalize `DeleteAt` onto `Threads`.)
Fixes: https://mattermost.atlassian.net/browse/MM-42919
We check for the presence of binary_parameters
in the DSN and add the 0x01 byte accordingly.
This helps us avoid casting to string
and efficiently use the database.
```release-note
NONE
```
Co-authored-by: Mattermod <mattermod@users.noreply.github.com>
The older method used to reply completely on timestamps
to take batches of items in a timestamp range and then
just incrementing the timestamp. This led to handling
edge-cases such as more items than the batch count, all
having the same timestamp.
Additionally, relying on timestamp as the page cursor
meant that indexing was not very efficient if you had
several items spread out across large spans of time.
To get away from all of that we use a proper cursor-based
approach consisting of createAt+Id. With this, we move
completely to a constant page size where we can fetch
a given number of objects irrespective of when they
were created. This makes indexing much more faster and
efficient.
https://mattermost.atlassian.net/browse/MM-41260
```release-note
Elasticsearch and Bleve indexing have been revamped to be much
more efficient and faster. The config parameter BulkIndexingTimeWindowSeconds
for both elasticsearch and bleve have been removed.
A new config parameter called BatchSize has been introduced instead.
This parameter controls the number of objects that
can be indexed in a single batch. This makes things
more efficient and maintains a constant workload.
```
We implement a cursor based pagination model
to page through the posts in a given thread.
The cursor is a combination of the post.CreateAt+
post.Id to differentiate multiple posts in a given
timestamp.
Some additional parameters like direction, fromPost,
fromCreateAt and perPage were introduced to implement
this.
```release-note
NONE
```
Due to the way our community deployment is done. The job server
is restarted every day. This means that whenever there is a job
that takes more than 24 hours, it will always get cancelled
when the server restarts and therefore will never finish.
This PR adds ability to resume any stopped jobs, by storing
intermediate progress in the job metadata and setting
the job to pending instead of cancelled when everything is
shut down.
The user can still cancel a job explicitly by clicking on
the cross button in the system console. That functionality
hasn't changed. Only server stop or stopping/starting
job server via config will pause/resume jobs.
```release-note
The elasticsearch indexing job is resumable now. Stopping a
server while the job is running will put the job in pending status
and will resume the job when the server starts.
The job can still be explicitly cancelled via the system console UI.
```
* revamp db version and add applied migrations endpoint
* replace old schema version with new
* add db version subcommand
* add to local api
* reflect review comments
* log errors
* remove setting the version from model.CurrentVersion
* fix a test
* use different field for schema version
* add build hash and current version to the support packet
* add tests
* update test to use new assets
* MM-42282: handle teamId parameter correctly
As per https://community-daily.mattermost.com/core/pl/ugs7ue6e4j8a7cgegk1bxje8to, `ThreadStore.GetThreadsForUser` accepts a `teamId` parameter, but incorrectly handles an empty value of `""` as looking only for channels with an empty `teamId` (aka DMs and GMs) instead of finding all channels and effectively ignoring the team property.
Fixes: https://mattermost.atlassian.net/browse/MM-42282
* break up getThreadsForUser, leverage errgroup
This change breaks up `GetThreadsForUser` in the `ThreadStore` into its constituent `GetTotalUnreadThreads`, `GetTotalThreads`, `GetTotalUnreadMentions`, and the original `GetThreadsForUser` but now solely returning the thread structures. Instead of a monolithic method at the store level, the application layer now handles calling bulk requests, leveraging `errgroup` for simpler parallelization.
This change brings with it a few benefits:
* Simpler code, including more idiomatic usage of squirrel
* Simpler SQL, joining tables only when configured conditions require same. (No performance benefit here, since an unused LEFT JOIN generally has no overhead.)
* Discrete Grafana metrics for each store method, giving us better insight into the performance characteristics in play.
* **Performance boost**: reduced overhead when clearing push notifications.
This last point is what prompted the re-re-reactoring in this PR. As I broke things up, I realized that `clearPushNotificationSync` only used the `TotalUnreadMentions`, but asked for the count of total threads and total unread threads. By exposing the discrete methods, this code path avoids two aggregate queries. We clear notifications when marking a thread as read, and when marking a channel with unread mentions as viewed, so I expect we'll see at least a modest boost to performance from simply not wasting these cycles anymore.
No performance improvements are expected from this PR for the general case of using `GetThreadsForUser` to populate the threads view.
* never discard errors from building queries
* no MustSql
Co-authored-by: Mattermod <mattermod@users.noreply.github.com>
* about to merge with master.
* merge from latest master.
* refactor squirrel function names to match original. clean up extra debugging.
* fix lint.
* reverted websocket_norace_test
* Fix issue of incorrect error being returned.
* made changes from annotated code.
* fix small lint issue.
* remove comment.
* cleaned up code some more.
* fix error with archived tests.
* Cleanup
```release-note
NONE
```
* more cleanup
```release-note
NONE
```
* address review comments
```release-note
NONE
```
* sq.Eq optimization
```release-note
NONE
```
Co-authored-by: Agniva De Sarker <agnivade@yahoo.co.in>