Repository navigation
FileSystemWatcher may cause problems in containers - inotify limits and incorrect error message #27272
Description
Activity
Just an FYI, our friends at Zeit figured out their server-side misconfiguration here https://github.com/zeit/now-examples/pull/61 but it would have been easier with the better error message.
It looks like this issue as it applies to corefx is just to improve the error message? That seems quite reasonable, and easy to fix.
Hello, I've just opened a code branch to upgrade to EF Core 2.2. My new code branch is failing all CI builds with the infamous "inotify" error:
System.IO.IOException : The configured user limit (1024) on the number of inotify instances has been reached.
We have about 3,500 unit tests, and it only fails at some point really far along in the tests. Build bots are on Ubuntu 18.04.
I've read the discussion above, and all I can say is that it sounds relevant to my problem, but I have no idea how to fix it. Help, please?
We have about 3,500 unit tests, and it only fails at some point really far along in the tests.
It sounds like either you're not disposing of all of the FileSystemWatchers you're creating, or something is causing tons of tests that create FileSystemWatchers to run concurrently. If your tests aren't themselves creating FileSystemWatchers, then it sounds like something in the environment is creating them, maybe something in EF Core 2.2, and it'd likely be worth an issue in the EF Core repo, assuming that's where they're coming from.
@stephentoub I'm not explicitly creating any FileSystemWatchers. And my tests are running sequentially, not concurrently.
Here's a sample stack trace of a failing test:---> System.IO.IOException: The configured user limit (1024) on the number of inotify instances has been reached. at System.IO.FileSystemWatcher.StartRaisingEvents() at System.IO.FileSystemWatcher.StartRaisingEventsIfNotDisposed() at System.IO.FileSystemWatcher.set_EnableRaisingEvents(Boolean value) at Microsoft.Extensions.FileProviders.Physical.PhysicalFilesWatcher.TryEnableFileSystemWatcher() at Microsoft.Extensions.FileProviders.Physical.PhysicalFilesWatcher.CreateFileChangeToken(String filter) at Microsoft.AspNetCore.Mvc.RazorPages.Internal.PageActionDescriptorChangeProvider.GetChangeToken() at Microsoft.Extensions.Primitives.ChangeToken.OnChange(Func`1 changeTokenProducer, Action changeTokenConsumer) at Microsoft.AspNetCore.Mvc.Infrastructure.DefaultActionDescriptorCollectionProvider..ctor(IEnumerable`1 actionDescriptorProviders, IEnumerable`1 actionDescriptorChangeProviders) at lambda_method(Closure , IBuildSession , IContext ) --- End of inner exception stack trace --- at lambda_method(Closure , IBuildSession , IContext ) at StructureMap.Building.BuildPlan.Build(IBuildSession session, IContext context) at StructureMap.Pipeline.LazyLifecycleObject`1.CreateValue() at StructureMap.SessionCache.GetObject(Type pluginType, Instance instance, ILifecycle lifecycle) at StructureMap.SessionCache.GetDefault(Type pluginType, IPipelineGraph pipelineGraph) at StructureMap.Container.GetInstance(Type pluginType) at Microsoft.Extensions.DependencyInjection.ServiceProviderServiceExtensions.GetRequiredService[T](IServiceProvider provider) at Microsoft.AspNetCore.Builder.MvcApplicationBuilderExtensions.UseMvc(IApplicationBuilder app, Action`1 configureRoutes) at TestingUtilities.Web.AspNetTestConfiguration.Configuration(IApplicationBuilder app) in <snip path>/src/TestingUtilities.Web/AspNetTestConfiguration.cs:line 48Line 48 of my
AspNetTestConfiguration.csis:public void Configuration(IApplicationBuilder app) { .... // line 48: app.UseMvc(routes => { routes.MapRoute("Default", "{controller}/{action}/{id?}", new { controller = "Home", action = "Index" }); }); }This, in turn, is being called by my integration test's base class, which creates a new
TestServerfor each test fixture. I am calling.Dispose()on theTestServerin theTearDown()method of the test fixture.Also, to clarify, it wasn't just EF Core I upgraded; I meant to say that I upgraded to .NET Core 2.2.
@shaulbehr, are you able to attach a debugger to the process when it's in one of these states? e.g. if you could attach lldb and use sos, you could use
dumpheap -type FileSystemWatcherto see what FSWs are hanging around, whether they're disposed or not, and hopefully if not why not (e.g. what's keeping them alive). If you're able to try .NET Core 3.0, you could also try out the new dotnet-dump tool (https://devblogs.microsoft.com/dotnet/introducing-diagnostics-improvements-in-net-core-3-0/), which should make it easy to collect a dump of the process that can then be analyzed similarly with the sos commands.@natemcmaster, @rynowak, I'm not sure who's responsible for this support used from MVC, but have you seen any issues related to PhysicalFilesWatcher instances not being disposed of in a timely manner?
I don't think we've seen issues with disposal, but our file watchers tend to live the lifetime of the app. I think it was the case for a while that we didn't dispose file watchers so that could cause bugs if the app was stopped and started repeatedly in the same process.
Note: that as of 2.2 we shouldn't be creating the filewatcher from that call stack anymore when the environment is set to
Production. This was a mitigation on our part for this scenario because of how many times we heard about these problems in containers./cc @pranavkm
Look at that call stack again, this makes a little more sense now if these are integration tests.
@shaulbehr are your tests creating a
TestServerfor each test?@rynowak Yes, each test fixture creates a new
TestServer. In addition, I have some test fixtures that have tests running in parallel, in which case I create a newTestServerfor each test. I did add code in theTearDownmethods to ensure that the TestServers are disposed, but this doesn't appear to have helped.Oho, here's something I just noticed. I added some code to ensure that my
IContainerobjects (fromStructureMap) are disposed, and now the stack trace going throughMicrosoft.AspNetCore.Builder.MvcApplicationBuilderExtensions.UseMvc()doesn't appear in my logs. Now I have a bunch of other iNotify errors, with the following stack trace:System.IO.IOException : The configured user limit (1024) on the number of inotify instances has been reached. Stack Trace: at System.IO.FileSystemWatcher.StartRaisingEvents() at System.IO.FileSystemWatcher.StartRaisingEventsIfNotDisposed() at System.IO.FileSystemWatcher.set_EnableRaisingEvents(Boolean value) at Microsoft.Extensions.FileProviders.Physical.PhysicalFilesWatcher.TryEnableFileSystemWatcher() at Microsoft.Extensions.FileProviders.Physical.PhysicalFilesWatcher.CreateFileChangeToken(String filter) at Microsoft.Extensions.Primitives.ChangeToken.OnChange(Func`1 changeTokenProducer, Action changeTokenConsumer) at Microsoft.Extensions.Configuration.FileConfigurationProvider..ctor(FileConfigurationSource source) at Microsoft.Extensions.Configuration.Json.JsonConfigurationSource.Build(IConfigurationBuilder builder) at Microsoft.Extensions.Configuration.ConfigurationBuilder.Build() at Microsoft.AspNetCore.Hosting.WebHostBuilder.BuildCommonServices(AggregateException& hostingStartupErrors) at Microsoft.AspNetCore.Hosting.WebHostBuilder.Build() at Microsoft.AspNetCore.TestHost.TestServer..ctor(IWebHostBuilder builder, IFeatureCollection featureCollection)@rynowak here's your smoking gun pointing at
TestServer.@stephentoub I'm really a rookie at Linux. If you can give me step-by-step instructions how to attach lldb and use sos and dumpheap, I'm probably up to that.
@rynowak, based on your question "are your tests creating a TestServer for each test?" and the answer of "Yes", it seems like you may have some insights here?
One option would be to try and limit the number of TestServer instances you create. That might or might not be feasible given your requirements. If it's possible, I would expect creating fewer servers to speed up your test execution as well.
Another thing you could try, would be to change how configuration is wired up and remove the file watching. The actual problem reported by that call stack is one that we've already fixed in 3.0 dotnet/extensions#928
13 remaining items
The fix for the other part of this this is dotnet/extensions#928.
I'm not convinced this is fixed.
Isn't the issue actually here in PhysicalFilesWatcher:
https://github.com/dotnet/runtime/blob/master/src/libraries/Microsoft.Extensions.FileProviders.Physical/src/PhysicalFilesWatcher.cs#L134This class attempts to respect
DOTNET_USE_POLLING_FILE_WATCHERas it is constructed withpollForChanges=truein that case (by PhysicalFileProvider), and proceeds to registerPollingFileChangeTokens for use instead of watching the filesystem ... howeverTryEnableFileSystemWatcheris called regardless of the value ofPollForChanges.For context, I'm still getting the error:
System.IO.IOException : The configured user limit (128) on the number of inotify instances has been reached, or the per-process limit on the number of open file descriptors has been reached.albeit due to a mistake in my test parallelisation, but still shouldn't be creating any file system watchers when running on the sdk linux docker image.
Re-opening for the Microsoft.Extensions issue...
- addeduntriagedNew issue has not been triaged by the area ownerNew issue has not been triaged by the area owner
on Jun 4, 2020 Will close again. #37664 can be used to track the Microsoft.Extensions issue.
I'm having a very similar issue to this debugging a docker compose proj with vs 2019.
https://stackoverflow.com/questions/63493884/system-io-ioexception-function-not-implemented-in-createhostbuilderargs-buil
The call stack is very similar but the error Message is "Function not implemented"
Happens every time I recompile the react js that changes file in ClientApp/build- ghost locked as resolved and limited conversation to collaborators
on Dec 15, 2020
MOVED FROM dotnet/aspnetcore#3475
Looking around the web I'm seeing years of issues with FileSystemWatcher saying "The configured user limit (n) on the number of inotify instances has been reached."
UPDATE: Looks like https://github.com/dotnet/corefx/blob/a10890f4ffe0fadf090c922578ba0e606ebdd16c/src/System.IO.FileSystem.Watcher/src/System/IO/FileSystemWatcher.Linux.cs#L371 will assume when inotify_add_watch fails with an ENOSPEC it must but an issue with inotify instances being out of range. In fact, ENOSPEC can also mean "the kernel failed to allocate a needed resource." We had no way to know it was anything other than "too many files open." The error message is misleading.
Phrased differently. There's two Error Cases and we throw a message that implies there's just One.
This is becoming more prevalent in container situations in constrained sandboxes. I'm trying to deploy https://github.com/shanselman/superzeit (just clone and "now --public" or run locally with docker) to Zeit.co and I'm hitting this regularly. I don't think I'm hitting a limit. I think Zeit (and others) are blocking the syscall.
I think there are two issues here:
1 We should return a different error message if inotify_add_watch fails, and then circuit break so that FileSystemWatcher doesn't prevent the app from starting. If we CAN startup without a watch successfully, we should.
2 It seems DOTNET_USE_POLLING_FILE_WATCHER=1 is used in dotnet-watch and the aspnet file providers but the base System.IO FileSystemWatcher class doesn't support DOTNET_USE_POLLING_FILE_WATCHER? We should probably be consistent.
If I change reloadOnChange: false in Program.cs to bypass the first watch that is set on AppSettings.json, I end up hitting it later when Razor/MVC sets up its FileWatchers.
We need at a minimum, to have DOTNET_USE_POLLING_FILE_WATCHER respected everywhere. Another idea would be for a way to have FileSystemWatcher "fail gracefully." We need to test on systems with
Related issues?
@stephentoub @natemcmaster @muratg @pranavkm