We would call SetStatusOnline in a goroutine before actually calling HubRegister.
This could cause the message not to be sent after all, because there's no guarantee
that Register would actually happen before it.
To fix it, we just change the order of things.
Co-authored-by: mattermod <mattermod@users.noreply.github.com>
After registering the conn in the hub, we proceeded to send a
direct message to the user. We had changed it to send the direct message
in the same hub goroutine that handles the registration. This was the correct
behavior and fixes chances of having panics due to sending to closed channels.
However, often fixing something unearths some deeper underlying bug. This was
such a case :)
The issue was that register channel had a buffer size of 1. And we were sending
a direct message after registration. In the code to send direct message, we
were checking if the user has been registered or not, and if not, then skip it.
Therefore, since the register channel buffer was 1, it could very well be that
the select case would pick up the direct message send case first - in which
case it would not have been registered, and therefore no hello message would be sent.
The fix is to unbuffer the register and unregister channels. There does not seem
to be a valid reason to make these buffered channels. They are meant to be
synchronous operations, because the code following them assumes that the user
has been registered.
While here, we also remove all the time.Sleeps before waiting on the Response channel
because they are not required at all. Waiting on a channel is already blocking.
Co-authored-by: Ben Schumacher <ben.schumacher@mattermost.com>
* MM-24987 Dont call a.GetSchemeRolesForChannel and instead load the scheme using the channel directly
* Trigger CI
* MM-24987 Add a test for concurrent patch to channel moderation
* MM-24987 Cleanup
* MM-24828 Take into account group synced state of team and channel when getting groups for mentions
* Update app/notification_test.go
* i18n-extract
* Revert "i18n-extract"
This reverts commit dcb0426b98afa4646c26870c0e3a1236f99fda17.
* Trigger CI
Co-authored-by: mattermod <mattermod@users.noreply.github.com>
* Disable read/search db replicas in TE/E0
* fixing tests
* Removing unnecesary text.
* Updating without-license read-replicas config before store initialization
* Reconnecting to database after remove read replicas
* MM-23935 extend session expiry on user activity
- if user types anything before a session expires the session will be extended to now + session length
- ensures new session expiries are not written to DB too frequently
- new session store func for updating session ExpiresAt
- session length defaults for mobile and web/ldap changed from 180 days to 30 days
* Fixing system messages about non-visible users
* Adding unit tests to verify the new behavior
* Regenerating app layers
Co-authored-by: mattermod <mattermod@users.noreply.github.com>
We invert the skipSend condition and outdent the remaining block
to make the code a bit more idiomatic.
While here, we also change the dropping message level from info to
warn because that's what it should be.
On user activity, we were clearing the job.pendingNotifications map.
But we had already created a copy of the notifications slice while
iterating the map. Therefore, if we pass the copied slice, it would
still have the old notifications which were originally deleted.
The unit tests would not catch this because it was testing the
job.pendingNotifications map and not actually checking if the email
handler was being called or not. We fix that now.
Co-authored-by: mattermod <mattermod@users.noreply.github.com>
* MM-23800: remove goroutineID and stack printing
Each hub has a goroutineID which is calculated with a known hack.
The FAQ clearly explains why goroutines don't have an id:
https://golang.org/doc/faq#no_goroutine_id.
We only added that because sometimes the hub would be deadlocked and
having the goroutineID would be useful when getting the stack trace.
This is also problematic in stress tests because the hubs would
frequently get overloaded and the logs would unnecessarily have stack traces.
But that was in the past, and we have done extensive testing with
load tests and fuzz testing to smooth any rough edges remaining.
Including adding additional metrics for hub buffer size.
Monitoring the metrics is a better way to approach this problem.
Therefore, we remove these kludges from the code.
* Also remove deadlock checking code
There is no need for that anymore since
we are getting rid of the stack printing anyways.
Let's do a wholesale refactor and clean up the codebase.
* MM-23805: Refactor web_hub
This is a beginning of the refactoring of the websocket code.
To start off with, we unexport some methods and constants which did not
need to be exported. There are more remaining but some are out of scope for this PR.
The main chunk of refactor is to unexport the webconn send channel
which was the main cause of panics. Since we were directly sending
to the connection from various parts of the codebase, it would be possible
that the send channel would be closed and we could still send a message.
This would crash the server.
To fix this, we refactor the code to centralize all sending from the main
hub goroutine. This means we can leverage the connections map to check
if the connection exists or not, and only then send the message.
We also move the cluster calls to cluster.go.
* bring back cluster code inside hub
* Incorporate review comments
* Address review comments
* rename index
* MM-23807: Refactor web_conn
- Unexport some struct fields and constants which are not necessary
to be accessed from outside the package. This will help us moving
the entire websocket handling code to a separate package later.
- Change some empty string checks to check for empty string rather
than doing a len check which is more idiomatic. Both of them compile
to the same code. So it doesn't make a difference performance-wise.
- Remove redundant ToJson calls to get the length.
- Incorporate review comments
- Unexport some more methods
* Fix field name
* Run make app-layers
* Add note on hub check
The current code path for `CreatePostAsUser` tries to update the `LastViewedAt` for bots, which logs a warning message if the bot isn't actually in the channel.
This pull request changes the semantics of bot posting to not update the bot's `LastViewedAt` timestamp, avoiding this log altogether. It matches the semantics of `from_webhook`, but notably makes bots slightly less like users in that they no longer "read" channels when they post. This seems reasonable, but I'm both looking for validation of this semantic change in addition to the code review.
Fixes: https://mattermost.atlassian.net/browse/MM-23926
Co-authored-by: mattermod <mattermod@users.noreply.github.com>
* MM-23800: remove goroutineID and stack printing
Each hub has a goroutineID which is calculated with a known hack.
The FAQ clearly explains why goroutines don't have an id:
https://golang.org/doc/faq#no_goroutine_id.
We only added that because sometimes the hub would be deadlocked and
having the goroutineID would be useful when getting the stack trace.
This is also problematic in stress tests because the hubs would
frequently get overloaded and the logs would unnecessarily have stack traces.
But that was in the past, and we have done extensive testing with
load tests and fuzz testing to smooth any rough edges remaining.
Including adding additional metrics for hub buffer size.
Monitoring the metrics is a better way to approach this problem.
Therefore, we remove these kludges from the code.
* Also remove deadlock checking code
There is no need for that anymore since
we are getting rid of the stack printing anyways.
Let's do a wholesale refactor and clean up the codebase.
Co-authored-by: mattermod <mattermod@users.noreply.github.com>
* MM-23017 Check group mentions as part of notification logic
* Add nil groups to existing test cases
* MM-23017 Add tests for insertGroupMention and addGroupMention
* MM-23017 Add tests for getExplicitMentions that have groups
* Add tests for group store GetMemberUsersNotInChannel
* MM-23017 Add tests for AllowGroupMentions
* MM-23017 Fix error message name
* MM-23017 Swap Checks to Name
* MM-23017 Code review fixes
* Rename var and fix allowGroupMentions test
* MM-23017 Use GetMemberUsersInTeam inside of insertGroupMentions
* MM-23017 use group mentions permission
* Actually call GetMemberUsersInTeam
* Remove unnecessary new line
* Uncomment filter allow reference
* MM-23017 Fix group channel notifications
* Update store layer
* MM-23017 Improve test coverage for group channels
* Trigger CI
* Trigger CI
* MM-20934: Fixing int overflow in 32 bits on MaxImageSize check
* Adding comments explaining the casting and the bug fixed there
* Apply suggestions from code review
Co-Authored-By: Juho Nurminen <juhonurm@gmail.com>
* Fixing store layers
Co-authored-by: Juho Nurminen <juhonurm@gmail.com>
* MM-23244: Validate that either both or neither AuthData and AuthService fields are set.
* MM-23244: Readability improvement.
* MM-23244: Adds translation. Tests for error id.
* MM-23244: Fix test.
Co-authored-by: mattermod <mattermod@users.noreply.github.com>
* Add database server version to telemetry
Also added a new query in the store to retrieve the database version
* Add test for the GetDbVersion function
* More drivers in the tests
Co-authored-by: mattermod <mattermod@users.noreply.github.com>
Add auditing to server CLI.
Also:
- simplify auditing in API layer
- reduce number of AddMeta calls
- have models serialize themselves
- more consistent field naming
* MM-23620: Handle error from GetUser
In case of high DB load, the DB will start to throw errors.
Unless we handle the error appropriately, the server will crash.
* Removing unnecessary lines
* MM-23770: Fix for blank DefaultChannelGuestRole on team schemes.
* MM-23770: Adds test for an team scheme with a blank DefaultChannelGuestRole field.
* MM-23770: Fix for unexpected Schemes.Get.
* store failed timestamps on health check job instead of on registeredPlugin
Update test
* change EnsurePlugin calls
* Make env.SetPluginState private
* Write test for plugin deactivate and PluginStateFailedToStayRunning
* Add license comment
* adjust comments, use time.Since
* Additional PR feedback:
time.Since cleanup
test cleanup
remove duplicate .Store() call
* PR Feedback
- Add test case for reactivating the failed plugin
- Change `crashed` to `healthy` and `hasPluginCrashed` to `isPluginHealthy`
- remove stale timestamps from health check job
* Keep registeredPlugins in env when plugin is deactivated, so the crashed state of a plugin can be persisted.
* PR feedback
* PR feedback from Jesse
Co-authored-by: mattermod <mattermod@users.noreply.github.com>