Skip to content

Troubleshooting / Container Cannot Run Normally

1. Container Keeps Restarting in a Docker Stack Environment

This problem is generally caused by incorrect configuration, firewall rules, and various whitelist settings.

Specifically, it manifests as:

  1. The page cannot be opened in a browser.
  2. When using the sudo docker ps -a command to view the container list, you may notice that the container keeps restarting.
  3. Running curl http://localhost:8088 locally on the deployment server returns the error curl: (7) Failed to connect to localhost port 8088: Connection refused.
  4. The log file continuously outputs error stack information.

Possible causes and solutions:

Possible Cause Solution
Manually modified configuration but the configuration contains errors Check the modified configuration files and verify whether the YAML syntax, database connection information, etc. are correct
The configuration specifies an external server, but the network is unreachable Check the firewall, Alibaba Cloud security group settings, database connection whitelist, and other configurations
Compatibility issues with the operating system See Compatibility Issues with the Operating System below
Redis does not support the current system's page size See Redis Does Not Support the Current System's Page Size below

Compatibility Issues with the Operating System

When using docker logs {Server Container ID}, if you encounter the following or similar error:

Text Only
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
node[9]: ../src/node_platform.cc:61:std::unique_ptr<long unsigned int> node::WorkerThreadsTaskRunner::DelayedTaskScheduler::Start(): Assertion `(0) == (uv_thread_create(t.get(), start_thread, this))' failed.
 1: 0xb57f90 node::Abort() [node]
 2: 0xb5800e  [node]
 3: 0xbc915e  [node]
 4: 0xbc9230 node::NodePlatform::NodePlatform(int, v8::TracingController*, v8::PageAllocator*) [node]
 5: 0xb1b3d1 node::InitializeOncePerProcess(int, char**, node::InitializationSettingsFlags, node::ProcessFlags::Flags) [node]
 6: 0xb1bc89 node::Start(int, char**) [node]
 7: 0x7f2ca389fd90  [/lib/x86_64-linux-gnu/libc.so.6]
 8: 0x7f2ca389fe40 __libc_start_main [/lib/x86_64-linux-gnu/libc.so.6]
 9: 0xa93f0e _start [node]
Aborted (core dumped)

This may be caused by incompatibility between the current operating system/components and Docker (for example, DataFlux Func 2.x bundles Docker 20.10.8, which may have issues on the latest versions of operating systems).

You can try the following methods to resolve it:

Upgrade DataFlux Func to the latest version. During the upgrade, please allow the installation Script to upgrade the Docker version as well. For details, see Deployment and Maintenance / Upgrade and Restart / Upgrade System

If the latest version of DataFlux Func still has the above issue, users can also download the official Docker binary package and upgrade to a newer version of Docker.

For downloads, please visit the official Docker download page: https://download.docker.com/linux/static/stable/

Or the Alibaba Cloud mirror site: https://mirrors.aliyun.com/docker-ce/linux/static/stable/

For example, on Ubuntu, you can use the following commands to upgrade operating system components:

Bash
1
2
3
sudo apt update
sudo apt upgrade
sudo apt dist-upgrade

Redis Does Not Support the Current System's Page Size

When the official Redis image starts on certain ARM-based operating systems, the error <jemalloc>: Unsupported system page size occurs. See:

2. Container Does Not Exist in a Docker Stack Environment

This problem is generally caused by an incorrect runtime environment.

Specifically, it manifests as:

  1. Run sudo docker stack ls and you can see dataflux-func.
  2. Run sudo docker ps -a and you cannot see the corresponding container.
  3. Run sudo docker stack ps dataflux-func --no-trunc and find that the container status is abnormal.

Possible causes and solutions:

Possible Cause Solution
Docker installed via Snap Uninstall the Snap version of Docker, reinstall the official Docker, or use the Docker bundled with the Script
Others You can troubleshoot based on the ERROR column in sudo docker stack ps dataflux-func --no-trunc

A typical example is no space left on device, which indicates insufficient disk space.

3. Container Cannot Start in a k8s Environment

This problem is generally caused by issues with the host machine / k8s cluster.

The following errors may occur in k8s:

Text Only
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
Events:
  Type     Reason   Age                   From     Message
  ----     ------   ----                  ----     -------
  Warning  Failed   36m                   kubelet  Error: failed to start container "func-server": Error response from daemon: OCI runtime create failed: container_linux.go:367: starting container process caused: process_linux.go:495: container init caused: rootfs_linux.go:60: mounting "/home/cce/kubelet/pods/xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx/volume-subpaths/user-config/func-server/1" to rootfs at "/home/cce/docker/overlay2/xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx/merged/data/user-config-template.yaml" caused: no such file or directory: unknown
  Warning  Failed   36m                   kubelet  Error: failed to start container "func-server": Error response from daemon: OCI runtime create failed: container_linux.go:367: starting container process caused: process_linux.go:495: container init caused: rootfs_linux.go:60: mounting "/home/cce/kubelet/pods/xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx/volume-subpaths/user-config/func-server/1" to rootfs at "/home/cce/docker/overlay2/xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx/merged/data/user-config-template.yaml" caused: no such file or directory: unknown
  Warning  Failed   36m                   kubelet  Error: failed to start container "func-server": Error response from daemon: OCI runtime create failed: container_linux.go:367: starting container process caused: process_linux.go:495: container init caused: rootfs_linux.go:60: mounting "/home/cce/kubelet/pods/xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx/volume-subpaths/user-config/func-server/1" to rootfs at "/home/cce/docker/overlay2/xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx/merged/data/user-config-template.yaml" caused: no such file or directory: unknown
  Normal   Created  35m (x5 over 118d)    kubelet  Created container func-server
  Warning  Failed   35m                   kubelet  Error: failed to start container "func-server": Error response from daemon: OCI runtime create failed: container_linux.go:367: starting container process caused: process_linux.go:495: container init caused: rootfs_linux.go:60: mounting "/home/cce/kubelet/pods/xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx/volume-subpaths/user-config/func-server/1" to rootfs at "/home/cce/docker/overlay2/xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx/merged/data/user-config-template.yaml" caused: no such file or directory: unknown
  Normal   Pulled   33m (x6 over 118d)    kubelet  Container image "dataflux-func.com/dataflux-func:2.7.0" already present on machine
  Warning  BackOff  2m8s (x157 over 36m)  kubelet  Back-off restarting failed container

The following errors may occur in a Func service:

Text Only
1
2
3
4
5
6
Traceback (most recent call last):
File "get-config.py", line 11, in <module>
CONFIG = yaml_resource.load_config(os.path.join(BASE_PATH, './config.yaml'))
File "/usr/src/app/worker/utils/yaml_resources.py", line 83, in load_config
user_config_content = _f.read()
OSError: [Errno 5] Input/output error

This is not a DataFlux Func issue. Please check the host machine / k8s cluster, and if NAS is involved, also check whether the NAS has any issues.