GitVP开源文摘
全部文章/算法面试

我的自学笔记

我的自学笔记,终身更新

作者huangruiteng 仓库huangruiteng/CS-Notes ↗ 星标★ 3,998 字数153,675 许可MIT 阅读1
GitHub 原文 ↗
摘要笔记按主题分层,文件都在 `Notes/` 下(同名子目录存放该笔记的图片与资源)。

CS-Notes

2021年版Intro

  • 我的自学笔记,在学习MLSys和C++,整理算法、操作系统,后续学习分布式系统,终身更新。
  • 我不是贵系oi出身的大神,目前在以一个初学者的心态补习计算机知识,学习路径和心得可能更值得大家借鉴。

2025年版Intro

  • 笔记的 AI 自动化管理,参考 .trae/documents 文件夹中的内容
  • 多了太多内容,无从谈起,建议看目录。

笔记目录

笔记按主题分层,文件都在 Notes/ 下(同名子目录存放该笔记的图片与资源)。

  • 计算机基础
    • 操作系统与系统
- [APUE](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/APUE.md) / [OSTEP](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/OSTEP-Operating-Systems-Three-Easy-Pieces.md)
- [Linux多线程服务端编程-muduo](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/Linux多线程服务端编程-muduo.md) / [操作系统](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/操作系统.md)
- [CSAPP](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/CSAPP.md) / [CSAPP-Labs](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/CSAPP-Labs.md) / [Shell-MIT-6-NULL](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/Shell-MIT-6-NULL.md)
- [Computer-Architecture](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/Computer-Architecture.md) / [Assembly](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/Assembly.md) / [Compiling](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/Compiling.md) / [Metaprogramming](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/Metaprogramming.md)
  • 网络与通信
- [通信与网络](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/通信与网络.md)
- [CS144-Lab](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/Computer-Networking-Lab-CS144-Stanford.md) / [CS144-Lecture](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/Computer-Networking-Lecture-CS144-Stanford.md)
- [Machine-Learning](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/Machine-Learning.md) / [AI-Algorithms](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/AI-Algorithms.md) / [AI-Applied-Algorithms](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/AI-Applied-Algorithms.md)
- [深度学习推荐系统](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/深度学习推荐系统.md) / [计算广告](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/计算广告.md)
- [Reinforcement-Learning](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/Reinforcement-Learning.md) / [Federated-Learning](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/Federated-Learning.md)
  • 大模型与系统
- [LLM-MLSys](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/LLM-MLSys.md) / [MLSys+RecSys](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/MLSys+RecSys.md)
- [GPU](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/GPU.md) / [pytorch](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/pytorch.md) / [tensorflow](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/tensorflow.md)
- [AI-Agent-Engineering](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/AI-Agent-Engineering.md) / [AI-Agent-Product&PE](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/AI-Agent-Product&PE.md)
- [Codex-Subagent:架构、执行、通信与恢复](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Notes/Codex-Subagent.md)

笔记心得

当前公开学习材料见 Learning-Materials:Top30 与 ranked backlog、其余候选,以及数据库与分布式系统补课讲义。

笔记与素材的管理代码、规则和测试集中在 note-system,与学习内容分开维护。

  • 用「子标题」清晰表达结构
  • 在「子标题」压缩信息:提炼正文内容

关于本仓库

  • This repository uses Github as image host, for the simplicity and reliability of backup.
  • 顶层目录中,Notes/ 放长期笔记,snippets/ 放可复用脚本与代码片段,prompts/ 放可复用 prompt 资产。
  • 本仓库内容由自制笔记转化器自动生成
    • remaining bug: 包含$的行会视为latex行内公式转化为公式图片,但这样会将shell代码转换,需要判断是否在latex代码域内,予以排除
    • solution: clone仓库,用Typora阅读Note文件夹里的本体文件
  • 内推字节AML团队,请联系 huangrt01@163.com

社交媒体

Star History

Star History Chart


OSTEP Operating Systems Three Easy Pieces

[toc]

__主题:virtualization, concurrency, persistence__

book, book-code, projects, homework xxxzz's homework answer, my homework answer

Intro

todo 南京大学副教授蒋炎岩:Mosaic操作系统模型和检查器_哔哩哔哩_bilibili

1.Dialogue

I hear and I forget. I see and I remember. I do and I understand. 其实是荀子说的

2.Introduction to Operating Systems

  • Von Neumann model
  • OS:并行,外设,resource manager
  • 概念:virtualization, API, system calls, standard library
CRUX: how to virtualize resources
CRUX: how to build correct concurrent programs
  • persistance
  • write有讲究:1)先延迟一会按batch操作 2)protocol, such as journaling or copy-on-write 3)复杂结构B-tree
CRUX: how to store data persistently
  • 目标:
    • performance: minimize the overheads
    • protection ~ isolation
    • reliability
    • security
    • mobility
    • energy-efficiency
  • history:
    • libraries
    • protection
    • trap handler user/kernel mode
    • multiprogramming minicomputer
    • memory protection concurrency ASIDE:UNIX
    • modern era: PC Linus Torvalds: Linux

Kernel

  • Meltdown patch
  • Meltdown patch 之后,Page Table Isolation (PTI),syscall 会 flush TLB
  • Linux 4.14 之后,support PCID (process-context id)
* Process-Context Identifiers (PCID) enables us to achieve the same goal of isolation much more efficiently by associating some data with each TLB entry which the processor uses to control access to the mappings. By changing the PCID during the mode switch, TLB entries with the kernel’s PCID will not be accessible from user space.

CPU性能优化 - KV Store

《Put an Elephant into a Fridge: Optimizing Cache Efficiency for In-memory Key-value Stores》

Abstract

现状: Such an extremely low cache-to-memory ratio (less than 0.1%) poses a significant new challenge—the limited CPU cache is becoming a severe performance bottleneck that hinders us from fully exploiting the great potential of high-speed memory-based key-value stores.

问题:cache contention, thrashing, inability to scale, https://en.wikipedia.org/wiki/Thrashing_(computer_science)

解决方案:By carefully reorganizing

the data layout in memory, redesigning the hash indexing structure, and offloading garbage collection, we can effectively improve the utilization of the limited cache space.

1.Introduction

1.1 Techinical Trend and Challenges

  • CPU cache 贵、重要、对性能影响大、scalable能力有限
  • 利用 CPU cache 的 challenges
  • hardware challenges: 软件很难直接管理cpu cache
  • software challenges: 数据冷热特点;value size比key大,caching a value可能逐出多个key;hash indexing structure很重要

1.2 Making Key-Value Store Cache Aware

software-only solution: best placement of a key-value item according to its temporal locality; key/value分离;新的hash indexing structure

  1. Motivations and Challenges
  • Issue #1: Disproportional key and value sizes
  • Issue #2: Low cache utilization in hash indexing.
  • Issue #3: Read amplification with key-values.

3.Mechanism

  • page coloring: 本质上给LLC cache做了32个分片
  • Gaining control on cache
  • Mapping with Hugepage: 2-MB page = 16 columns (one column = 32 rows 64 sets 64 B)
  • Mapping with pre-allocated pages
  • get_pgcolor

4.Policy

4.1 Handling Hot and Cold Key-value Data

  • we desire to see that each row is filled up with both cold and hot data, which compete for the space within the row, and upon eviction, the victims would be the cold ones.
  • 手段:Memcached maintains an LRU list per slab class to track each key-value item’s relatively locality, while Redis maintains a pool of weak-locality (cold) key-values for eviction by sampling the dataset periodically.

4.2 Separating Key and Value Data

  • placing them in separate rows
  • key/value 读写分离的问题:
  • parallel access: 维护两个指针,缺点是需要指针以及内存带宽的浪费
  • concurrent access: 维护两个queue,效果更好

4.3 Cache-friendly Hash Indexing

case study 1: Upon inserting a key-value item, a slab from the slab class with the smallest slot size that can accommodate the item is selected.

首先,尽量让hot items(logical position接近的items)集中在某列(physical zone)

其次,每一个physical zone内部的hot item尽量不在同一个row(不同的color)

  1. Case Study 1: MEMCACHED
  • 6.1 Optimizations
    • head-to-head allocation比较有趣,在key-value分离的设计下,解决了不知道key and value内存比例的问题
  1. Case Study 2: REDIS
  • redis和memcached的区别:没有slab -> 不需要 LRU list 来做 repartition,用一个LRU clock,后台线程做eviction

8.Discussion

  • 只要cashe sets够用,增加set数量不会有提升
  • huge page由于是system-wide配置,有受到恶意攻击disturb the shared cache的可能

MICA: A Holistic Approach to Fast In-Memory Key-Value Storage, NSDI 2014

https://www.usenix.org/conference/nsdi14/technical-sessions/presentation/lim

MICA(Memory-store with Intelligent Concurrent Access) takes a holistic approach that encompasses all aspects of request handling, including parallel data access, network request handling, and data structure design, but makes unconventional choices in each of the three do- mains. First, MICA optimizes for multi-core architectures by enabling parallel access to partitioned data. Second, for efficient parallel data access, MICA maps client requests directly to specific CPU cores at the server NIC level by using client-supplied information and adopts a light-weight networking stack that bypasses the kernel. Finally, MICA’s new data structures—circular logs, lossy concurrent hash indexes, and bulk chaining—handle both read- and write-intensive workloads at low overhead.

  1. Key Design Choices

3.1 Parallel Data Access

MICA’s parallel data access: MICA partitions data and mainly uses exclusive access to the partitions. MICA exploits CPU caches and packet burst I/O to disproportionately speed more loaded partitions, nearly eliminating the penalty from skewed workloads. MICA can fall back to concurrent reads if the load is extremely skewed, but avoids concurrent writes, which are always slower than exclusive writes. Section 4.1 describes our data access models and partitioning scheme.

3.2 Network Stack

  • socket I/O 比较费,大量的read
  • direct NIC access
  • request direction
  • Flow-level core affinity: 1) Receive-Side Scaling (RSS); 2) Flow Director (FDir)
  • Object-level core affinity
  • MICA's request direction: 利用Flow Director,在client去做object的编码,让NIC能理解

3.3 KV Data Structures

3.3.1 Memory Allocator

  • cache mode: log structure
  • store mode: segregated fits

3.3.2 Indexing: Read-oriented vs. Write-friendly

  • lossy data structures:
  • bulk chaining
  • use memory allocator's eviction support to avoid evicting recently-used items (4.3.2)
  1. MICA Design
4.1.2 分析CREW(Concurrent Read Exclusive Write)模式 -> 4.3.1

4.2 Network Stack

​ UDP 4.2.1 Direct NIC Access

​ Intel's DPRK

4.2.2 Client-Assisted Hardware Request Direction

4.3 Data Structure

MICA, in cache mode, uses circular logs to manage memory for key-value items and lossy concurrent hash indexes to index the stored items. Both data structures exploit cache semantics to provide fast writes and simple memory management. Each MICA partition consists of a single circular log and lossy concurrent hash index. MICA provides a store mode with straightforward extensions using segregated fits to allocate memory for key- value items and bulk chaining to convert the lossy concurrent hash indexes into lossless ones.

Hugepages (2 MiB in x86-64) use fewer TLB entries for the same amount of memory, which significantly reduces TLB misses during request processing.

hugepage必须预先一次分配2M或者1GB的内存空间,并且使用mmap的接口去分配。所以用malloc和free的库来管理大页是合理的(比如jemalloc去支持hugepage配置)。

4.3.2 Lossy Concurrent Hash Index

MEMC3: Compact and concurrent memcache with dumber caching and smarter hashing, NSDI 2013

  1. Optimistic Concurrent Cuckoo Hashing
  • An optimistic version of cuckoo hashing that supports multiple-reader/single writer concurrent access, while preserving its space benefits;
    • 记录version,version不对(比如是奇数)就retry
    • 利用了 the atomic read/write for 64-bit aligned pointers on 64-bit machines 的特点,参考 APPENDIX
  • A technique using a short summary of each key to improve the cache locality of hash table operations;
    • An optimization for cuckoo hashing insertion that improves throughput
    • CLOCK-based approximate LRU,clock算法可以增强thread的scale能力(消除LRU synchronization的瓶颈)

Full-stack architecting to achieve a billion-requests-per-second throughput on a single key-value store server platform, 2016

network-with-kv

杂项

内存优化

  • garbage collection
    • 《Quantifying the performance of garbage collection vs.
explicit memory management.》

Virtualization

3.Dialogue

4.the abstraction: The Process —— 线程状态

### S(Sleeping,可中断睡眠状态 ) - 含义:线程处于休眠状态,此时线程在等待某个事件发生,比如等待 I/O 操作完成、等待信号量、等待其他线程释放资源等。处于该状态的线程不会占用 CPU 时间,只有当等待的事件发生时,线程才会被唤醒并进入可运行状态。 - 举例:当一个线程发起了一个网络请求后,在等待服务器响应的过程中,它就会进入可中断睡眠状态,直到接收到网络数据后才会被唤醒继续执行。 ### R(Running 或 Runnable,运行或可运行状态 ) - 含义 - Running:线程正在被 CPU 执行,此时线程占用 CPU 资源,正在执行指令。 - Runnable:线程已经准备好执行,在就绪队列中等待 CPU 调度,一旦获取到 CPU 资源,就会立即进入 Running 状态开始执行。 - 举例:一个计算密集型的线程,在执行复杂的数学运算时,就处于 Running 状态;而多个这种计算线程同时存在时,除了正在被 CPU 执行的那个线程,其他等待被调度执行的线程就处于 Runnable 状态。 ### D(Disk sleep,不可中断睡眠状态 ) - 含义:线程处于深度睡眠状态,通常是在等待 I/O 操作(如磁盘 I/O )完成。与可中断睡眠状态不同,处于不可中断睡眠状态的线程不能被信号中断,只有当相关的 I/O 操作完成后,线程才会被唤醒。这是为了保证 I/O 操作的原子性和完整性,防止在 I/O 操作过程中线程被意外中断而导致数据不一致等问题。 - 举例:当线程执行文件读取操作时,如果此时磁盘繁忙,线程需要等待磁盘数据准备好,它就会进入不可中断睡眠状态,直到磁盘将数据准备好传输给线程。 ### Z(Zombie,僵死状态 ) - 含义:线程已经终止运行,但是其父进程还没有调用相应的函数(如 wait () )来收集其终止信息,回收其资源。处于僵死状态的线程虽然已经不再运行,但仍在进程表中占据一个位置,保留一些资源信息,直到父进程处理。 - 举例:在父子进程模型中,子进程先于父进程结束运行,但父进程没有及时调用 wait () 函数获取子进程的退出状态,此时子进程就会处于僵死状态。 ### T(Traced 或 Stopped,跟踪或停止状态 ) - 含义 - Traced:线程被调试器(如 gdb )跟踪,处于被调试的状态,此时线程的执行会受到调试器的控制,比如可以暂停、单步执行等。 - Stopped:线程被信号(如 SIGSTOP )暂停执行,处于停止状态,直到收到继续执行的信号(如 SIGCONT )才会恢复运行。 - 举例:当使用 gdb 调试一个多线程程序时,被调试的线程就处于 Traced 状态;当在终端中向某个线程发送 SIGSTOP 信号时,该线程就会进入 Stopped 状态。
CRUX: how to provide the illusion of many CPUs
  • low level machinery
    • e.g. context switch : register context
  • policies
    • high level intelligence
    • e.g. scheduling policy
  • separating policy(which) and mechanism(how)
    • modularity
  • Process Creation
    • load lazily: paging and swaping
    • run-time stack; heap(malloc(),free())
    • I/O setups; default file descriptors
进程状态转移
  • final态(在UNIX称作zombie state)等待子进程return 0,parent进程 wait()子进程
# 查找僵尸进程
ps -aux|grep Z
ps -ef|grep 子进程pid
kill -9 父进程pid
  • xv6 process structure
// the registers xv6 will save and restore
// to stop and subsequently restart a process
struct context {
    int eip;int esp;
    int ebx;int ecx;
    int edx;int esi;
    int edi;int ebp;
};
// the different states a process can be in
enum proc_state { UNUSED, EMBRYO, SLEEPING,RUNNABLE, RUNNING, ZOMBIE };
// the information xv6 tracks about each process
// including its register context and state
struct proc {
    char*mem;                  // Start of process memory
    uint sz;                    // Size of process memory
    char*kstack;               // Bottom of kernel stack
                               // for this process
    enum proc_state state;      // Process state
    int pid;                    // Process ID
    struct proc*parent;        // Parent process
    void*chan;                 // If !zero, sleeping on chan
    int killed;                 // If !zero, has been killed
    struct file*ofile[NOFILE]; // Open files
    struct inode*cwd;          // Current directory
    struct context context;     // Switch here to run process
    struct trapframe*tf;       // Trap frame for the
                                // current interrupt
};
  • Data Structure: process list,PCB(Process Control Block)
  • HW:process-run.py
    • -I IO_RUN_IMMEDIATE 发生IO的进程接下来会有IO的概率大,所以这样高效

5.Interlude: Process API

CRUX: how to create and control processes
  • #include ,getpid(),fork() 不从开头开始运行
  • scheduler的non-determinism,影响concurrency
  • p3.c 利用execvp执行子程序wc
    • reinitialize the executable,transform原进程
    • 不会return
    • exec调用会把当前进程的机器指令都清除,因此前后的printf都不会执行
  • fork+exec的意义: it lets the shell run code after the call to fork() but before the call to exec(); this code can alter the environment of the about-to-be-run program, and thus enables a variety of interesting features to be readily built.
  • p4.c
#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>
#include <string.h>
#include <fcntl.h>
#include <assert.h>
#include <sys/wait.h>

int main(int argc, char *argv[])
{
    int rc = fork();
  //  printf("STDOUT_FILENO的值是%d",STDOUT_FILENO);
    if (rc < 0) {
        // fork failed; exit
        fprintf(stderr, "fork failed\n");
        exit(1);
    } else if (rc == 0) {
	// child: redirect standard output to a file

	close(STDOUT_FILENO); 
	open("./p4.output", O_CREAT|O_WRONLY|O_TRUNC, S_IRWXU);

	// now exec "wc"...
        char *myargs[3];
        myargs[0] = strdup("wc");   // program: "wc" (word count)
        myargs[1] = strdup("p4.c"); // argument: file to count
        myargs[2] = NULL;           // marks end of array
        execvp(myargs[0], myargs);  // runs word count
    } else {
        // parent goes down this path (original process)
        int wc = wait(NULL);
	assert(wc >= 0);
    }
    return 0;
}
  • file descriptor的原理:按序搜索,因此需要close(STDOUT_FILENO);
  • 类似的应用:UNIX的pipe()特性,grep -o foo file | wc -l
  • 谁可以发送SIGINT信号给process=>signal(), process group, 引入user的概念
  • RTFM:read the fucking manual

HW:

fork()和vfork()的区别:

  1. fork ():子进程拷贝父进程的数据段,代码段
* vfork( ):子进程与父进程共享数据段
  1. fork ()父子进程的执行次序不确定
* vfork 保证子进程先运行,在调用exec 或exit之前与父进程数据是共享的,在它调用exec或exit 之后父进程才可能被调度运行。
  1. vfork ()保证子进程先运行,在她调用exec 或exit 之后父进程才可能被调度运行。如果在调用这两个函数之前子进程依赖于父进程的进一步动作,则会导致死锁。
  • 5.4 不同的exec#C_language_prototypes)
    • execvp,p的含义是寻找路径,v:vector
  • 5.5 如果child没有child,在child里用wait没有意义
  • 5.6 waitpid() wait和waitpid的区别
    • The pid parameter specifies the set of child processes for which to wait. If pid is -1, the call waits for any child process. If pid is 0, the call waits for any child process in the process group of the caller. If pid is greater than zero, the call waits for the process with process id pid. If pid is less than -1, the call waits for any process whose process group id equals the absolute value of pid.
  • 5.8 注意子进程返回0
管道
  • stdio buffering
    • It should be noted here that changing the buffering for a stream can have unexpected effects. For example glibc (2.3.5 at least) will do a read(blksize) after every fseek() if buffering is on.
    • 期望环境变量控制
* `tail -f access.log | BUF_1_=1 cut -d' ' -f1 | uniq`
* 风险:Denial Of Service possibilities
  • fcntl(fileno(stdin), F_SETPIPE_SZ, pipe_page_num * 4096);
    • 16MB is system max value by default for socket
import fcntl
import platform

try:
    if platform.system() == 'Linux':
        fcntl.F_SETPIPE_SZ = 1031
        fcntl.fcntl(fd, fcntl.F_SETPIPE_SZ, size)
except IOError:
    print('can not change PIPE buffer size')

6.Mechanism: Limited Direct Execution

CRUX: how to efficiently virtualize the cpu with control
  • limited direct execution
CRUX: how to perform restricted operations
  • aside: open() read()这些系统调用是trap call,写好了汇编,参数和系统调用number都放入well-known locations
    • 概念:trap into the kernel return-from-trap trap table trap handler
    • be wary of user inputs in secure systems
  • NOTE:
    1. x86用per-process kernel stack,用于存进程的寄存器值,以便trap的时候寄存器够
    2. 如何控制:set up trap table at boot time;直接进任何内核地址是very bad idea
    3. user mode不能I/O request
  • system call,包括accessing the file system, creating and destroying processes, communicating with other processes, and allocating more memory(POSIX standard)
* protection: user code中存在的是system call number,避开内核地址
* 告诉硬件trap table在哪也是privileged operation
LDE protocal

stub code

CRUX: how to regain control of the CPU
  • problem #2:switching between processes
    • A cooperative approach: wait for system calls
    • MacOS9 Emulator
    • NOTE: only solution to infinite loops is to reboot the machine,reboot is useful。
    • 重启的意义:1)回到确定无误的初始状态;2)回收过期和泄漏的资源;3)不仅适合手动操作,也易于自动化
    • A Non-Cooperative Approach: The OS Takes Control
CRUX: how to gain control without cooperation
  • a timer interrupt interrupt handler
    • timer也可以关
  • deal with malfeasance: in modern systems, the way the OS tries to handle such malfeasance is to simply terminate the offender.
  • scheduler context switch
LDE protocal + timer interrupt

注意有两种register saves/restores:

  • timer interrupt: 用hardware,kernel stack,implicitly,存user registers
  • OS switch:用software,process structure,explicitly,存kernel registers
  • e.g. xv6 context switch code

NOTE:

  • 如何测time switch的成本:LMbench
  • 为何这么多年操作系统速度没有明显变快:memory bandwidth
  • 如何处理concurrency?=> locking schemes,disable interrupts
    • 思考:baby-proof

HW: measurement

7.Scheduling: Introduction

CRUX: how to develop scheduling policy
  • workload assumptions
    • fully-operational scheduling discipline
    • 概念:jobs
  • scheduling metrics:turnaround time

FIFO: convoy effect $\stackrel{\bf{允许长度不等}}{\longrightarrow}$ SJF(shortest job first) $\stackrel{\bf{允许来时不等}}{\longrightarrow} $ STCF(Shortest Time-to-Completion First )=PSJF $\stackrel{\bf{允许不run\ to\ completion}}{\longrightarrow}$ a new metric: response time

  • 概念:preemptive schedulers
  • Round-Robin(RR) scheduling(轮转调度算法)
    • time slice scheduling quantum
    • 时间片长:amortize the cost of context switching
  • 针对I/O:overlap

8.Scheduling: The Multi-Level Feedback Queue(MLFQ) 多级反馈队列

  • Corbato图灵奖;和security有联系
CRUX: How to schedule without perfect knowledge

多个queue,每个queue对应一个priority,队内用RR => how to change priority

  • Rule 1:If Priority(A)>Priority(B), A runs (B doesn’t).
  • Rule 2:If Priority(A)=Priority(B), A & B run in RR.

attempt1: how to change priority

  • Rule 3:When a job enters the system, it is placed at the highest priority (the topmost queue).
  • Rule 4a:If a job uses up an entire time slice while running, its priority is reduced(i.e., it moves down one queue).
  • Rule 4b:If a job gives up the CPU before the time slice is up, it stays at the same priority level.
  • 思考:是否优先级越高的queue越倾向于用RR

MLFQ的问题:

  1. starvation
  2. game the scheduler
  3. change its behavior

attempt2: the priority boost

  • Rule 5:After some time period S, move all the jobs in the system to the topmost queue.
    • 部分地解决1和3
    • 和思考一致:高优先级,把time slice调短
    • Solaris:Default values for the table are 60 queues, with slowly increasing time-slice lengths from 20 milliseconds (highest priority) to a few hundred milliseconds (lowest),and priorities boosted around every 1 second or so,

attempt3: better accounting

  • Rule 4:Once a job uses up its time allotment at a given level (regardless of how many times it has given up the CPU), its priority is reduced (i.e., it moves down one queue).

其它可能的特性:

  • 操作系统0优先级
  • advice机制:As the operating system rarely knows what is best for each and every process of the system, it is often useful to provide interfaces to allow usersor administrators to provide some hints to the OS. We often call such hints advice, as the OS need not necessarily pay attention to it, but rather might take the advice into account in order to make a better decision. Such hints are useful in many parts of the OS, including the scheduler(e.g., with nice), memory manager (e.g.,madvise), and file system (e.g.,informed prefetching and caching [P+95])
  • HW: iobump,io结束后把进程调到当前队列第一位,否则最后一位;io越多效果越好

9.Scheduling: Proportional share

应用: C++ 进程/线程优先级

实时和非实时调度策略测试总结

  • Linux内核的三种调度策略:
    • SCHED_OTHER 分时调度策略
    • SCHED_FIFO实时调度策略,先到先服务。一旦占用cpu则一直运行。一直运行直到有更高优先级任务到达或自己放弃
    • SCHED_RR实时调度策略,时间片轮转。当进程的时间片用完,系统将重新分配时间片,并置于就绪队列尾。放在队列尾保证了所有具有相同优先级的RR任务的调度公平
    • pthread_setschedparam(id, policy, &param)
CRUX: how to share the CPU proportionally

Basic Concept: Tickets Represent Your Share

  • 利用randomness:
    1. 避免corner case,LRU replacement policy (cyclic sequential)
    2. lightweight
    3. fast,越快越伪随机

NOTE:

如果对伪随机数限定范围,不要用rand,用interval

机制:

  1. ticket currency,用户之间
  2. ticket transfer,用户与服务器
  3. ticket inflation,临时增加tickets,需要进程之间的信任
  • unfairness metric
  • stride scheduling — deterministic
    • lottery scheduling 相对于 stride 的优势:no global state
curr = remove_min(queue);   // pick client with min pass
schedule(curr);             // run for quantum
curr->pass += curr->stride; // update pass using stride
insert(queue, curr);        // return curr to queue

9.7 The Linux Completely Fair Scheduler(CFS) 完全公平调度器

  • 引入 vruntime的概念,记录进程运行时间
  • 引入sched_latency=48ms, time slice=48/n,保证均分
    • min_granularity=6ms,防止n太大sched_latency过低的情况
    • Weighting (Niceness) --->time slice ; table对数性质,ratio一致
  • 用红黑树储存进程节点
    • sleep,则从树中移去
    • 关于I/O,回来后设成树里的最小值

NOTE:

  • 这个idea应用广泛,比如用于虚拟机的资源分配
  • why index-0?

10.Multiprocessor Scheduling (Advanced)

  • 概念:multicore processor 与threads配合
CRUX: how to schedule jobs on multiple CPUs

Background: Multiprocessor Architecture

单核与多核的区别:the use of hardware caches(e.g., Figure 10.1), and exactly how data is shared across multiple processors.

Q: cache coherence问题,不同CPU的cache未及时更新,导致读错数据

A: 利用hardware, by monitoring memory accesses: bus snooping; write-back caches

Don’t Forget Synchronization

虽然解决了coherence问题,依然需要mutual exclusion primitives(locks)

One Final Issue: Cache Affinity

Single-Queue Scheduling

SQMS (single-queue multiprocessor scheduling),存在的问题如下:

  • lack of scalability: lock overhead随CPU数目增加而变高
  • cache affinity: 需要用复杂的机制调度,比如将少量jobs migrating from CPU to CPU

->

MQMS (multi-queue multiprocessor scheduling): 分配jobs给CPU

  • 优点:more scalable; intrinsically provides cache affinity
  • 缺点:load imbalance
CRUX: how to deal with load imbalance?

migration: 将余数migrate

the tricky part: how should the system decide to enact such a migration?

e.g. work stealing: source queue经常peek其它的target queue,使两边work load均衡

  • peek频率是超参数,过高会失去MQMS的意义,过低会有load imbalances
Linux Scheduling
  • Linux Multiprocessor Schedulers
  • O(1) scheduler: priority-based scheduler
  • CFS (Completely Fair Scheduler): a deterministic proportional-share approach
  • BFS (BF Scheduler): also proportional-share, but based on a more complicated scheme known as Earliest Eligible Virtual Deadline First (EEVDF)
  • BFS是single queue,前两个是multiple queues
  • Linux中,每个CPU都有一个自己的本地队列(红黑树),用来存放等待这个CPU资源的,已就绪待运行的tasks。正在运行的task,主要有两种机制离开CPU。
    • 自愿抢占:通过调用sleep, lock, io等主动让出CPU;
    • 非自愿抢占: 被优先级更高的task抢占或者时间片到期,被动的离开CPU。
  • 显然,task如果在CPU队列上等待的时间较长,一定会影响请求延迟。 但事实上Linux的CPU调度算法是非常优秀的,而且CPU是一种可以抢占的资源,因此通常即使在一个比较高的CPU利用率下,延迟也不会急剧上升(和磁盘IO调度相比)。

fcfdaf2f-b76b-47d9-a49f-ca74513cb09f

11.summary

12.A dialogue on memory virtualization

every address generated by a user program is a virtual address

  • ease of use, isolation, protection

13.The Abstraction: Address Spaces

  • multiprogramming
  • abstraction: address space
CRUX:how to virtualize memory
  • virtual memory
    • goals:transparency, efficiency (e.g. TLBs), protection
  • location of code : 0x105f40ec0
  • location of heap : 0x105f55000
  • location of stack: 0x7ffee9cbf8ac
  • 64bit系统下进程的内存分布
    • Linux64位操作系统仅使用低47位,高17位做扩展(只能是全0或全1)。所以,实际用到的地址为256TiB,空间为0x0000000000000000 ~ 0x00007FFFFFFFFFFF(user space)和0xFFFF800000000000 ~ 0xFFFFFFFFFFFFFFFF(kernel space),其余的都是unused space。
    • 对于 32-bit Linux,一个进程的地址空间是 4GiB,其中用户态能访问 3GiB 左右, 而一个线程的默认栈(stack)大小是 10MB,心算可知,一个进程大约最多能同时启动 300 个线程

NOTE:

  • 用microkernels的思想实现isolation,机制和策略的分离

14.Interlude: Memory API

CRUX:how to allocate and manage memory
  • 64bit UNIX系统,int和double都是8个字节
  • automatic memory management ~ garbage collector
  • 其它calls:calloc()先置0,realloc()更大区域

一些常见错误:

  • segmentation fault =>用strdup ###
  • buffer overflow e.g. 应该strlen(src)+1
  • uninitialized read/undefined value ###
  • memory leak 针对long-running server,是OS层面的错误
  • dangling pointer
  • double free ###
  • invalid frees
  • strcat的参数内存区域重复 ###
  • (###:valgrind可检测的)
  • 用purify和valgrind检查内存泄漏
  • 底层基础:

HW:

  • null.c segmentation fault
  • forget_free.c lldb没反应
  • dangling_pointer.c 直接run会print出0
  • free_wrong.c int *类型的+操作符重载过,直接加数字,不用乘sizeof(XXX)

15.Mechanism: Address Translation

efficiency and control

CRUX: how to efficiently and flexibly virtualize memory

hardware-based address translation

dynamic (hardware-based) relocation=base and bounds

  • physical address = virtual address + base
  • base and bounds register,本质上是MMU(memory management unit)
  • bounds register: 两种方式,一是size,二是physical address

<->

static (software-based) relocation: loader,不安全,难以再次换位

一些硬件要素:寄存器,异常,内核态,privileged instructions

  • 硬件和protection联系紧密

OS需要的数据结构:

  • free list(定长进程内存)
  • PCB(or process structure) 储存base和bounds信息
  • exception handlers: 掐掉过界的进程
LDE+Dynamic Relocation

问题:internal fragmentation,内存利用率不高 => segmentation

Dynamic Relocation

16.Segmentation

CRUX:how to support a large address space
  • 意义:节省内存,不赋予全部虚拟地址空间以实体
  • 概念:sparse address spaces、segmentation、segmentation fault
  • 实现
    • explicit:利用前两位;也可只利用一位,把code和heap合并
    • implicit:利用计组知识,比如PC生成的地址属于code区
    • 对于stack的特殊处理:negative offset

support for sharing

  • code sharing 这是一个潜在的好处
  • protection bits (硬件支持)
  • fine-grained segmentation: segment table
  • coarse-grained

OS support

  • segment registers
  • malloc
  • manage free space
    • 问题:external fragmentation
    • 方案1: compaction
* 消耗大
* makes requests to grow existing segments hard to serve
  • 方案2:free-list:best-fit,worst-fit,first-fit(address-based ordering,利于coalesce),buddy algorithm (块链表,合并)
  • What is memory fragmentation?
    • virtual memory下的memory fragmentation问题没那么严重
    • 自己维护内存池可能还是有必要的,相当于利用了业务先验知识(某些内存块是一起分配一起释放的、以及能感知locality)
    • It's when you have mixtures of short-lived and long-lived objects that you're most at risk, but even then malloc will do its best to help.
  • How to solve Memory Fragmentation
    • perf free blocks in the heap的情况
    • logically divide the heap into sub arenas where the lifetimes are more similar
    • to attempt to make the allocation sizes more similar or identical
* **unnecessary** because the allocator will already be doing a power of two size bucketing scheme

17.Free-Space Management

CRUX: how to manage free space
  • 重点是external fragmentation
  • memory分配给user后,禁止compaction

Low-level mechanisms: 运用了以下机制

  • Splitting and Coalescing
  • Tracking the size of allocated regions
  • void free(void*ptr) {header_t*hptr = (header_t*) ptr - 1;}
  • size 和 magic;寻找 chunk 的时候需要加上头的大小
  • Embedding A Free List
    • typedef struct __node_t {int size; struct __node_t *next;} node_t;
  • Growing The Heap

Other Approaches: (这些 approaches 的问题在于 lack of scaling)

  • segregated list,针对高频的size
    • slab allocator: 利用了这个特性,object caches,pre-initialized state
  • binary buddy allocator

并行优化:Hoard,jemalloc (radix trees 存 metadata)

18.Paging: Introduction

另一条路径,page frames,never allocating memory in variable-sized chunks

CRUX: how to virtualize memory with pages
  • flexibility and simplicity
  • page table: store address translations
    • inverted page table不是per-process结构,用哈希页表,可以配合TLB

translate: virtual address= virtual page number(VPN) + offset

=> PFN(PPN): physical frame(page) number

18.2 Where are page tables stored?

  • page table大,仅用hardware MMU难以管理,作为virtualized OS memory,存在physical memory里,甚至可以存进swap space

PTE: page table entry

  • valid bit:x86的实现中,没有valid bit,由OS利用额外的结构决定,present bit=0的page,是否valid,即是否需要swapped back in
  • protection bits
  • present bit (e.g. swapped out)
  • dirty bit
  • reference bit (accessed bit) ~ page replacement
  • 实际存储:VirtualAddress:32=20(VPN)+12(Offset);PTE内部:20(PFN)+3(empty)+9(flag)
accessing memory with paging

19.Paging: Faster Translations(TLBs)

CRUX: how to speed up address translation

TLB: translation-lookaside buffer

  • 属于MMU,是address-translation cache
  • TLB hit/miss
  • cache:spatial and temporal locality ; locality 是一种 heuristic
  • TLB是全相联cache

TLB Control Flow Algorithm

VPN = (VirtualAddress & VPN_MASK) >> SHIFT
(Success, TlbEntry) = TLB_Lookup(VPN) 
if(Success == True) // TLB Hit
    if (CanAccess(TlbEntry.ProtectBits) == True)
        Offset   = VirtualAddress & OFFSET_MASK
        PhysAddr = (TlbEntry.PFN << SHIFT) | Offset
        Register = AccessMemory(PhysAddr)
    else
        RaiseException(PROTECTION_FAULT)
else                  // TLB Miss
    PTEAddr = PTBR + (VPN*sizeof(PTE))
    PTE = AccessMemory(PTEAddr)
    if (PTE.Valid == False)
        RaiseException(SEGMENTATION_FAULT)
    else 
    		if (CanAccess(PTE.ProtectBits) == False)
        		RaiseException(PROTECTION_FAULT)
        else
            TLB_Insert(VPN, PTE.PFN, PTE.ProtectBits)
            RetryInstruction()

OS-handled,实现细节:

  • 普通的return-from-trap回到下条指令,TLB miss handler会retry,回到本条指令
  • 防止无限循环:trap handler放进physical memory,或者对部分entries设置wired translations
VPN = (VirtualAddress & VPN_MASK) >> SHIFT
(Success, TlbEntry) = TLB_Lookup(VPN)
if (Success == True)   // TLB Hit
    if (CanAccess(TlbEntry.ProtectBits) == True)
        Offset   = VirtualAddress & OFFSET_MASK
        PhysAddr = (TlbEntry.PFN << SHIFT) | Offset
        Register = AccessMemory(PhysAddr)
    else
        RaiseException(PROTECTION_FAULT)
else                  // TLB Miss
    RaiseException(TLB_MISS)

19.3 Who Handles The TLB Miss?

  • Aside: RISC vs CISC
  • CISC: x86: hardware-managed
    • multi-level page table
    • 硬件知道PTBR
    • the current page table is pointed to by the CR3 register [I09]
  • RISC: MIPS: software-managed

ASIDE: TLB Valid Bit和Page Table Valid Bit的区别:

  1. PTE和新进程密切相联。
  2. context switch时把TLB valid bit置0

19.5 TLB Issue: Context Switches

CRUX: how to manage TLB contents on a context switch
  • solution1: flush the TLB
    • 对于硬件实现,PTBR的变化后flush the TLB
  • solution2: ASID(address space identifier) 8bit ,性质上类似于32bit的PID
  • NOTE: 可能存在进程间的 sharing pages,比如库或者代码段

Issue: cache replacement policy

CRUX: how to design TLB replacement policy
  • LRU, 会有corner-case behaviors
  • random policy
  • MIPS 的 TLBs 是 software-managed,一个 entry 64bit,有 32或64个 entry,会给OS预留,比如用于 TLB miss handler
A MIPS TLB Entry
  • MIPS TLB 相关的四个 privileged OS 命令:TLBP(probe), TLBR(read), TLBWI(replace specific), TLBWR(replace random)
  • Culler's Law:TLB经常是性能瓶颈
  • Issue: exceeding the TLB converge => larger pages,应用于DBMS
  • Issue: physically-indexed cache成为bottleneck => virtually-indexed cache [W03]
    • in the CPU pipeline, with such a cache, address translation has to take place before the cache is accessed(计组知识)
  • HW:测量NUMPAGES,UNIX: getpagesize()=4096

20.Paging: Smaller Tables

CRUX: how to make page tables smaller?

bigger pages, multiple page sizes, DBMS

hybrid approach: paging and segments

  • hybrid的思想,尤其针对看似对立的机制
  • e.g. Multics
  • base: physical address of the page table , limit:how many valid pages
  • 有三个page tables=>三个base registers而不是一个
  • issue:
    1. page table的大小可变,与内存相联系,重新产生了external segmentation
    2. 不灵活,比如不针对堆很稀疏的情形
SN = (VirtualAddress & SEG_MASK) >> SN_SHIFT
VPN = (VirtualAddress & VPN_MASK) >> VPN_SHIFT
AddressOfPTE = Base[SN] + (VPN*sizeof(PTE))

multi-level page tables

  • page directory 好处:加入了level of indirection,更灵活 坏处:Time-space trade-off
  • PDE(page directory entry)
  • 如何分组:page size/PTE size ~ n位 n位一组即可

inverted page tables

  • 本质上是个数据结构问题
  • size ~ 物理页数 < 进程数*虚拟页数
  • swapping the page tables to disk: VAX/VMS

21.Beyond Physical Memory: Mechanisms

CRUX: how to go beyond physical memory
  • hard disk drive
  • 与single address space对立的旧机制:memory overlays
  • swap space:在硬盘上,disk address
  • the present bit
    • 0的意义:page fault handler
    • 为什么称作fault?是因为硬件无法处理,需要raise an exception交给OS

page replacement policy

  • Page-Fault Control Flow Algorithm (Hardware)
VPN = (VirtualAddress & VPN_MASK) >> SHIFT
(Success, TlbEntry) = TLB_Lookup(VPN) 
if(Success == True) // TLB Hit
    if (CanAccess(TlbEntry.ProtectBits) == True)
        Offset   = VirtualAddress & OFFSET_MASK
        PhysAddr = (TlbEntry.PFN << SHIFT) | Offset
        Register =AccessMemory(PhysAddr)
    else
        RaiseException(PROTECTION_FAULT)
else                  // TLB Miss
    PTEAddr = PTBR + (VPN*sizeof(PTE))
    PTE =AccessMemory(PTEAddr)
    if (PTE.Valid == False)
        RaiseException(SEGMENTATION_FAULT)
    else 
        if (CanAccess(PTE.ProtectBits) == False)
            RaiseException(PROTECTION_FAULT)
        else if (PTE.Present == True)// assuming hardware-managed TLB
            TLB_Insert(VPN, PTE.PFN, PTE.ProtectBits)
            RetryInstruction()
        else if (PTE.Present == False)
            RaiseException(PAGE_FAULT)
  • Page-Fault Control Flow Algorithm (Software)
PFN = FindFreePhysicalPage()
if(PFN == -1) // no free page found
    PFN = EvictPage()       // run replacement algorithm
DiskRead(PTE.DiskAddr, PFN)// sleep (waiting for I/O)
PTE.present = True          // update page table with present
PTE.PFN     = PFN           // bit and translation (PFN)
RetryInstruction()          // retry instruction 

when replacements really occurs

  • swap(page) daemon的任务:可用页数低于LW(low watermark)时free到HW
    • cluster a number of pages
  • daemon:守护进程
    • 可以救活coreaudiod这种进程
  • idle time: background,比如把文件写入memory而非disk
  • 总结:以上这些,对process是透明的

HW:

vmstat命令

  • vmstat 1 显示每秒状态
  • 运行多个,user time变大,idle time 变少
  • 运行1024MB,swpd不变,free减少,exit之后还原
  • cat /proc/meminfo 可用内存132GB,运行巨量mem会core dumped,in(中断时间)明显增加,偶尔会有sy(system time)
  • swapon -s,显示可供swap的大小 ,相当于 cat /proc/swaps

22.Beyond Physical Memories: Policies

CRUX: how to decide which page to evict
  • 本质上是cache management
  • 评价指标average memory access time:$ AMAT = T_{M}+(P_{Miss} · T_{D}) $,hit rate很重要

the optimal replacement policy

  • Belady: furtherest in the future

ASIDE: types of cache misses: Three C’s: compulsory, capacity, conflict

  • OS page cache是全相联,不会发生conflict miss

A simple policy: FIFO

  • Belady’s Anomaly: FIFO,cache size高可能反而不好,因为没有stack property(N+1-cache和N-cache的包含关系),不像LRU有这个性质

another simple policy : random

Using History: LRU least-recently-used

  • frequency and recency LFU
  • 其它变种:scan resistance

workload有几种:随机,80-20(适合LRU),looping sequential(适合RAND)

CRUX: how to implement an LRU replacement policy

关于实现:approximating LRU

  • use bit(reference bit)
  • clock algorithm: 循环数组,遇到1置为0,遇到0置换
    • 效果比其它的好,只比LRU差一点
    • 改进:考虑dirty bits,先evict unused and clean pages,否则swap损耗大

其它policy:

  • page selection: demand paging/prefetching
  • clustering(grouping) of writes

thrashing: memory is oversubscribed, demands > physical memory

  • 方法一:admission control,控制working sets的大小
  • 方法二:out-of-memory killer, 潜在的问题:kill X server

23.Complete Virtual Memory Systems

CRUX: how to build a complete VM system
VAX/VMS virtual memory

DEC发明

存在的问题:

  • 需要覆盖的机器类型range太宽,the curse of generality
  • 有inherent flaws

page+segmentation

Q:page太小,512bytes,如何解决内存压力?

  1. hybrid approach: 引入segmentation
  2. 利用内核memory
  3. 与2联系,利用TLB缓解复杂机制带来的损耗
具体实现

NOTE:

  • page 0: in order to provide some support for detecting null-pointer accesses
  • kernel is mapped into each address space

有关page replacement:

  • no reference bit
  • 针对memory hog:segmented FIFO,给每个进程规定一个RSS(resident set size),超出这个范围要FIFO
  • second-chance lists: clean-page list和dirty-page list,从clean开始evict
  • clustering

other neat tricks:

  • demand zeroing: 等到进程要用page,再交给OS给page置0
  • COW: copy-on-write,和UNIX的fork() exec()机制结合
  • 核心思想是be lazy: 好处一是提高responsiveness,二是可能obviate the need to do things at all
the Linux virtual memory system

内核、用户部分的内存分配,大体沿用以前,区别在于内核分为两部分:

  • kernel logical addresses:
    • kmalloc,存page tables, per-process kernel stacks,不能swap到disk
    • direct mapping to physical addresses : 1) simple to translate,0xC0000000变0x00000000;2) contiguous,适合于DMA
  • kernel virtual addresses: vmalloc,不连续,用于large buffers,可以让32bit系统处理超过1GB的memory
  • 0xC0000000开始是内核

64-bit x86:

64-bit x86

large page support

  • explicit的支持:mmap, shmget => transparent huge page support
  • 对 TLB 好处大,miss rate 和 path 都降低成本
  • 缺点:internal fragmentation; swapping效果不好
  • 体现了 incrementalism,慢慢引入特性并迭代

the page cache

  • 主要来源:memory-mapped files, file data and metadata from devices , and heap and stack pages that comprise each process (anonymous memory)
  • pdflush:背景线程,把dirty data写入backing store
  • 2Q replacement:inactive list不时加入active list,解决大文件频繁访问的问题
  • memory mapping 体现在 linux 的方方面面: 用 pmap 命令查看

security相关

buffer overflow

smashing the stack for fun and profit

  • 概念:privilege escalation
  • 对策:NX bit page禁止执行,针对stack

return-oriented programming (ROP)

  • 对策:address space layout randomization(ASLR). even KALSR
    • macOS上,调整[27:12]16位,随机

Other Security Problems: Meltdown And Spectre

  • meltdownattack.com spectreattack.com
  • 思想:speculation execution会暴露内存信息
  • 对策:kernel page-table isolation (KPTI)

《cloud atlas》quote: “My life amounts to no more than one drop in a limitless ocean. Yet what is any ocean, but a multitude of drops?”

24.Summary

Concurrency

25.A Dialogue on Concurrency

26.Concurrency: An Introduction

  • Amdahl's Law: 1/(1-p)
    • 95% work ~ infinite concurrency, the theoretical speedup is limited to at most 20 times the single thread performance

概念:thread, multi-threaded, thread control blocks (TCBs)

  • thread-local: 栈不共用,在进程的栈区域开辟多块栈,不是递归的话影响不大
* [关于线程栈和进程栈](https://www.cnblogs.com/luosongchao/p/3680312.html)
  * 线程栈是固定大小的(默认8KB),可以使用`ulimit -a` 查看,使用`ulimit -s` 修改
  * 进程栈大小时执行时确定的,与编译链接无关
  * 进程栈大小是随机确认的,至少比线程栈要大,但不会超过2倍
  • thread的意义:1) parallelism, 2) 适应于I/O阻塞系统、缺页中断(需要KLT),这一点类似于multiprogramming的思想,在server-based applications中应用广泛。
015

NOTE:

  • pthread_join与detach
  • disassembler: objdump -d -g main
  • x86,变长指令,1-11个字节

问题:线程之间会出现data race,e.g. counter的例子

  • 经典面试题:两个线程,全局变量i++各运行100次,问运行完i的最小值。
    • 答案是2
进程2取 i=0
进程1执行99次
进程2算 i+1=0+1=1
进程1取 i=1
进程2执行99次
进程1算 i+1=1+1=2

引入概念:critical section, race condition, indeterminate, mutual exclusion

=> the wish for atomicity

  • transaction: the grouping of many actions into a single atomic action
  • 和数据库, journaling、copy-on-write联系紧密
  • 条件变量:用来等待而非上锁
CRUX: how to support synchronization
  • the OS was the first concurrent program!
  • Not surprisingly, pagetables, process lists, file system structures, and virtually every kernel data structure has to be carefully accessed, with the proper synchronization primitives, to work correctly.

HW26:

  • data race来源于线程保存的寄存器和stack
  • 验证了忙等待的低效

27.Interlude: Thread API

CRUX: how to create and control threads
#include <pthread.h>
int pthread_create(pthread_t*thread,const pthread_attr_t*attr,void*(*start_routine)(void*),void*arg);

typedef struct
{
    int a;
    int b;
} myarg_t;
typedef struct
{
    int x;
    int y;
} myret_t;
void * mythread(void *arg)
{
    myret_t *rvals = Malloc(sizeof(myret_t));
    rvals->x = 1;
    rvals->y = 2;
    return(void *)rvals;
}
int main(int argc, char *argv[])
{
    pthread_t p;
    myret_t * rvals;
    myarg_t args = {10, 20};
    Pthread_create(&p, NULL, mythread, &args);
    Pthread_join(p, (void **)&rvals);
    printf("returned %d %d\n", rvals->x, rvals->y);
    free(rvals);
    return 0;
}
  • pthread_create
    • thread: &p
    • attr:传参NULL或,pthread_attr_init
    • arg和start_routine的定义保持一致;
    • void=any type
  • pthread_join
    • simpler argument passing:(void *)100, (void **)rvalue
    • (void **)value_ptr,小心局部变量存在栈中,回传指针报错
  • gcc -o main main.c -Wall -pthread
lock
pthread_mutex_t lock;
pthread_mutex_lock(&lock);
x = x + 1; // or whatever your critical section is
pthread_mutex_unlock(&lock);
  • pthread_mutex_t lock = PTHREAD_MUTEX_INITIALIZER;
  • int rc = pthread_mutex_init(&lock, NULL); assert(rc == 0); // always check success!
  • pthread_mutex_trylock和timedlock
conditional variables
pthread_mutex_t lock = PTHREAD_MUTEX_INITIALIZER;
pthread_cond_t  cond = PTHREAD_COND_INITIALIZER;
Pthread_mutex_lock(&lock);
while (ready == 0) Pthread_cond_wait(&cond, &lock);
Pthread_mutex_unlock(&lock);

HW:

  • main-race.c:
    • valgrind --tool=helgrind ./main-race,结果给出了“Possible data race during write of size 4 at 0x30A014 by thread #1”
    • 全局变量存放在数据段
  • 误判了main-deadlock-global-c,说明有瑕疵
  • main-signal-cv.c 条件变量的用法示例
#include <stdio.h>
#include "mythreads.h"
// 
// simple synchronizer: allows one thread to wait for another
// structure "synchronizer_t" has all the needed data
// methods are:
//   init (called by one thread)
//   wait (to wait for a thread)
//   done (to indicate thread is done)
// 
typedef struct __synchronizer_t {
    pthread_mutex_t lock;
    pthread_cond_t cond;
    int done;
} synchronizer_t;

synchronizer_t s;

void signal_init(synchronizer_t *s) {
    Pthread_mutex_init(&s->lock, NULL);
    Pthread_cond_init(&s->cond, NULL);
    s->done = 0;
}

void signal_done(synchronizer_t *s) {
    Pthread_mutex_lock(&s->lock);
    s->done = 1;
    Pthread_cond_signal(&s->cond);
    Pthread_mutex_unlock(&s->lock);
}

void signal_wait(synchronizer_t *s) {
    Pthread_mutex_lock(&s->lock);
    while (s->done == 0) Pthread_cond_wait(&s->cond, &s->lock);
    Pthread_mutex_unlock(&s->lock);
}

void* worker(void* arg) {
    printf("this should print first\n");
    signal_done(&s);
    return NULL;
}

int main(int argc, char *argv[]) {
    pthread_t p;
    signal_init(&s);
    Pthread_create(&p, NULL, worker, NULL);
    signal_wait(&s);
    printf("this should print last\n");

    return 0;
}

28.Locks

  • 概念:lock, owner,
  • locks: 利用OS来schedule线程,套在critical section两边,有available(unlocked, free)和acquired(locked, held)两种状态
  • POSIX的mutex
CRUX: how to build a lock

evaluating locks

  • correctness (mutual exclusion), fairness (starve), performance
  • 需要考虑starve,设不设置contending
  • 多核问题
  • 定义锁的精细化程度:coarse和fine-grained lock

解决互斥问题应遵循的条件

  1. 任何两个进程不能同时处于临界区
  2. 不应对CPU的速度和数量做任何假设
  3. 临界区外运行的进程不得阻塞其他进程
  4. 不得使进程无限期等待进入临界区
禁止中断 interrupt masking
  • 优点:简单
  • 把禁止中断的权利交给用户进程导致系统可靠性较差
  • 不适用于多处理器(违反条件2)
  • 可能错过其它interrupts,比如disk的read request
  • 对于现代CPU,这个实现速度慢
  • 应用场景:OS内部数据结构的互斥访问
共享锁变量 just using loads/stores
while(lock==1);
lock=1;
//critical region
lock=0;
//non_critical region;
  • 违反条件1(interleaving)
  • 忙等待(spin-waiting)
  • Murphy's law
Building Working Spin Locks with Test-And-Set

while (TestAndSet(&lock->flag, 1) == 1);

  • test-and-set(atomic exchange):把返回原值+修改值这两个操作绑定
    • xchg(x86), ldstub(SPARC)
    • 可以test-and-test-and-set,只有当flag为0才改变锁。
  • spin lock,要求preemptive scheduler,抢占式调度
    • 满足correctness
    • 不满足fairness和performance
    • 在多核处理器上表现良好,因为当前的线程可以很快通过critical section,不需要多次spin,上下文切换
DEKKER’S AND PETERSON’S ALGORITHMS
#define FALSE 0
#define TRUE  1
#define N     2                                           /* 进程数量 */
  
int turn=0;                                 /* 现在轮到谁?*/
int interested[N];               /* 所有值初始化为0(FALSE)*/

void enter_region(int process)               /* 进程是0或1 */
{
	interested[process] = TRUE;              /* 表名所感兴趣的*/
	turn = 1-process;                            /* 设置标志 */
	while(turn == 1-process && interested[1-process] ==TRUE); /* 空语句 */
}

void leave_region(int process)                    
{
	interested[process] = FALSE;             /* 表示离开临界区*/ 
}
Compare-And-Swap
Load-Linked and Store-Conditional
typedef struct __lock_t{
	int flag;
} lock_t;

void init(lock_t *lock){
	//0:lock is available, 1:lock is held
	lock->flag = 0;
}

int LoadLinked(int *ptr)
{
	return *ptr;
}
int StoreConditional(int*ptr, int value) {
	if (no update to *ptr since LoadLinked to this address) {
		*ptr = value;
		return 1; // success!
	} 
	else {
		return 0; // failed to update
	}
}
void lock(lock_t *lock){
	while (LoadLinked(&lock->flag) ||!StoreConditional(&lock->flag, 1));
}
void unlock(lock_t *lock){
	lock->flag = 0;
}
Fetch-And-Add
  • ticket lock
    • 优点:ensure progress for all threads, 线程一定会被调度到
typedef struct __lock_t {
    int ticket;
    int turn;
} lock_t;
void lock_init(lock_t*lock) {
    lock->ticket = 0;
    lock->turn   = 0;
}
void lock(lock_t*lock) {
    int myturn = FetchAndAdd(&lock->ticket);
    while (lock->turn != myturn); // spin
}
void unlock(lock_t*lock) {
    lock->turn = lock->turn + 1;
}
CRUX: how to avoid spinning
方法一:just yield
  • 效率问题没解决,并且仍然有starvation问题
void lock() {
    while (TestAndSet(&flag, 1) == 1)
        yield(); // give up the CPU, running->ready
}
方法二:Using Queues: Sleeping Instead Of Spinning
  • OS support: park(), unpark() (Solaris)
  • 利用guard,虽然也有一定的spin lock损耗,但不涉及critical section,损耗较小
  • Q1:wakeup/waiting race:在park之前切换上下文
  • A1: 1)利用setpark(); 2)guard传入内核,可能类似后面futex的实现
  • Q2:priority inversion: 高优先级线程waiting低优先级线程,可能因为spin lock或者存在中优先级线程而无法运行。
  • A2: 1)priority inheritance; 2)所有线程平等;3)Priority ceiling protocol;4)Random boosting;5)Avoid blocking
方法三:Linux的futex,更多内核特性
  • OS support: per-futex in-kernel queue
  • nptl库中lowlevellock.h的代码片段:
    • mutex的巧妙设计
void mutex_lock (int*mutex) {
    int v;
    /*Bit 31 was clear, we got the mutex (the fastpath)*/
    if (atomic_bit_test_set (mutex, 31) == 0)
        return;
    atomic_increment (mutex);
    while (1) {
        if (atomic_bit_test_set (mutex, 31) == 0) {
            atomic_decrement (mutex);
            return;
        }
        /*We have to wait
        First make sure the futex value 
        we are monitoring is truly negative (locked).*/
        v =*mutex;
        if (v >= 0)
            continue;
        futex_wait (mutex, v);
    }
}
void mutex_unlock (int*mutex) {
    /*Adding 0x80000000 to counter results in 0 if and
    only if there are not other interested threads*/
    if (atomic_add_zero (mutex, 0x80000000))
        return;
    /*There are other threads waiting for this mutex,
    wake one of them up.*/
    futex_wake (mutex);
}
方法四:Two-Phase Locks
  • 在futex之前spin不止一次,可以spin in a loop
  • 思考:这又是一个hybrid approach(上一个是paging and segments)

29.Lock-based Concurrent Data Structures

CRUX: how to add locks to data structures

Concurrent Counters

  • 概念:thread safe, perfect scaling

Scalable Counting: approximate counter

  • local counter和global counter,一个CPU配一个锁,再加上一个global锁
  • threshold S: scalable的程度,local到global的transfer间隔
  • 实现见书上本章P5

LWN上的文章 :

  • atomic_t变量, SMP-safe (SMP:Symmetrical Multi-Processing),缺点在于锁操作的损耗、cache line频繁在CPU之间跳动
  • approximate counter:缺点在于耗内存、低于实际值
  • 进一步引入 local\_t,每个GPU设两个counter

Concurrent Linked Lists

  • malloc error之后接unlock,这样的代码风格容易出问题。实际实现时推荐只在update数据结构的时候加锁,因为malloc具有thread safe特性。
  • Tip:be wary of control flow changes that lead to function returns, exits, or other similar error conditions that halt the execution of a function
  • hand-over-hand locking(lock coupling):并发性强,但锁操作频繁,实际性能不见得好

Concurrent Queue

  • Michael and Scott Concurrent Queue: 1)头尾两个锁;2)头节点法
  • A more fully developed bounded queue, that enables a thread to wait if the queue is either empty or overly full, is the subject of our intense study in the next chapter on condition variables.

Concurrent Hash 基于concurrent lists

#define BUCKETS (101)
typedef struct __hash_t {
    list_t lists[BUCKETS];
} hash_t;
void Hash_Init(hash_t*H) {
    int i;
    for (i = 0; i < BUCKETS; i++)
        List_Init(&H->lists[i]);
}
int Hash_Insert(hash_t*H, int key) {
    return List_Insert(&H->lists[key % BUCKETS], key);
}
int Hash_Lookup(hash_t*H, int key) {
    return List_Lookup(&H->lists[key % BUCKETS], key);
}

premature optimization (Knuth's Law)

  • linux、Sun OS这样成熟的OS,为了规避这一问题,也是一开始只用 big kernel lock(BKL),等多核瓶颈出现后再做优化。 《Understanding the Linux Kernel 》

30.Condition Variables

CRUX: how to wait for a condition
  • 概念:condition variable, wait/signal on the condition
  • wait(): unlock, 然后让线程睡眠
  • signal(): lock,返回caller
int done  = 0;
pthread_mutex_t m = PTHREAD_MUTEX_INITIALIZER;
pthread_cond_t c  = PTHREAD_COND_INITIALIZER;
void thr_exit() {
    Pthread_mutex_lock(&m);
    done = 1;
    Pthread_cond_signal(&c);
    Pthread_mutex_unlock(&m);
}
void*child(void*arg) {
    printf("child\n");
    thr_exit();
    return NULL;
}
void thr_join() {
    Pthread_mutex_lock(&m);
    while (done == 0)
        Pthread_cond_wait(&c, &m);
    Pthread_mutex_unlock(&m);
}
int main(int argc, char*argv[]) {
    printf("parent: begin\n");
    pthread_t p;
    Pthread_create(&p, NULL, child, NULL);
    thr_join();
    printf("parent: end\n");
    return 0;
}
  • 代码中变量done的意义:state varibale,使程序的正确性不受两个线程运行先后顺序影响
  • 关于条件变量需要互斥量保护的问题,pthread_cond_wait内部先解锁再等待,之所以加锁是防止cond_wait内部解锁后时间片用完。https://blog.csdn.net/zrf2112/article/details/52287915
The Producer/Consumer (Bounded Buffer) Problem
  • bounded buffer的应用场景:HTTP requests的work queue;pipe

实现一:用一个条件变量+if实现

  • 问题:Mesa semantics: there is no guarantee that when the woken thread runs, the state will still be as desired <-> Hoare semantics;前者广泛采用

实现二:一中的if改成while,尽量不被遗漏

  • 多线程程序尽量用while来check条件,可以避免if条件满足时一次唤醒多个线程,spurious wakeups,资源不足
  • 问题:signal不确定唤醒的是生产者还是消费者

实现三:while+两个条件变量

Covering Conditions:指需要唤醒过多的满足条件的线程的情形

  • eg1: 针对memory allocator问题,直接用pthread_cond_broadcast唤醒所有wait中的线程,这是最简洁有效的思路
  • eg2: 生产者消费者问题的实现一,也有这个问题,但可以从原理上进行改进,而eg1不方便进行原理上的改进,只能broadcast

HW:

2../main-two-cvs-while -l 10 -m 10 -p 1 -c 1 -v -t -C 0,0,0,0,0,0,1

3.Linux switchs more often between producer and consumer than Mac

4.5.改m之后,由12/13秒到7秒

6.7.均是5秒,因为睡眠的时候释放了锁

9../main-one-cv-while -l 100 -p 1 -c 2 -m 1 -v -t

31.Semaphores

CRUX: how to use semaphores?

信号量和锁/条件变量的互相转换问题

#include <semaphore.h>
sem_t s;
sem_init(&s, 0, 1);
// second arg set to 0:the semaphore is shared between threads in the same process
// third arg: initial value

sem_wait(&m);
//critical section
sem_post(&m);
  • 信号量初始化的值如何选取:consider the number of resources you are willing to give away immediately after initialization
  • 信号量为负值时,绝对值是正在等待的线程数
  • sem_post(&m)运行不停滞

信号量的应用

  • Binary Semaphores
  • Semaphores For Ordering
  • The Producer/Consumer (Bounded Buffer) Problem
    • mutual exclusion
    • mutex在内层,否则会deadlock
void*producer(void*arg) {
	int i;
	for (i = 0; i < loops; i++) {
		sem_wait(&empty);       // Line P1
		sem_wait(&mutex);       // Line P1.5 (MUTEX HERE)
		put(i);                 // Line P2
		sem_post(&mutex);       // Line P2.5 (AND HERE)
		sem_post(&full);        // Line P3
	}
}

void*consumer(void*arg) {
	int i;
	for (i = 0; i < loops; i++) {
		sem_wait(&full);        // Line C1
		sem_wait(&mutex);       // Line C1.5 (MUTEX HERE)
		int tmp = get();        // Line C2
		sem_post(&mutex);       // Line C2.5 (AND HERE)
		sem_post(&empty);       // Line C3
		printf("%d\n", tmp);
	}
}
  • Reader-Writer Locks
    • 如果要保证公平竞争:设置互斥信号量S,加在读者/写者的acquire_lock函数上
    • 引申到设计理念,复杂往往低效,可能简单的spin lock更好;比如说CPU的cache设计,全相联比组相连效率高,部分是因为全相联实现的lookups更快
typedef struct _rwlock_t {
	sem_t lock;      // binary semaphore (basic lock)
	sem_t writelock; // allow ONE writer/MANY readers
	int   readers;   // #readers in critical section
} rwlock_t;

void rwlock_init(rwlock_t *rw) {
	rw->readers = 0;
	sem_init(&rw->lock, 0, 1);
	sem_init(&rw->writelock, 0, 1);
}

void rwlock_acquire_readlock(rwlock_t *rw) {
	sem_wait(&rw->lock);
	rw->readers++;
	if (rw->readers == 1) // first reader gets writelock
		sem_wait(&rw->writelock);
	sem_post(&rw->lock);
}

void rwlock_release_readlock(rwlock_t *rw) {
	sem_wait(&rw->lock);
	rw->readers--;
	if (rw->readers == 0) // last reader lets it go
		sem_post(&rw->writelock);
	sem_post(&rw->lock);
}

void rwlock_acquire_writelock(rwlock_t *rw) {
	sem_wait(&rw->writelock);
}
void rwlock_release_writelock(rwlock_t *rw) {
	sem_post(&rw->writelock);
}
  • The Dining Philosophers
    • 需要解决的问题:死锁,哲学家同时拿到左手的餐具,资源依赖成环
    • 方法一:对每个fork设置信号量;如下面代码所示,修改其中一位哲学家的get_forks()避免成环
    • 方法二:对每个人设置信号量,定义test函数,在test的外围加互斥锁
#define NUM 5
while(1){
	think();
	get_forks();
	eat();
	put_forks();
}

int left(int p) {return p;}
int right(int p) {return (p+1)%NUM;}

void put_forks(int p){
	sem_post(&forks[left(p)]);
	sem_post(&forks[right(p)]);
}
void get_forks(int p){
	if(p==NUM){
		sem_wait(&forks[right(p)]);
		sem_wait(&forks[left(p)]);
	} else{
		sem_wait(&forks[left(p)]);
		sem_wait(&forks[right(p)]);
	}
}
  • thread throttling

admission control, 比如针对memory-intensive region,避免thrashing(swap pages)

  • 如何实现信号量?
    • 用条件变量和互斥锁
    • 用信号量实现条件变量很难,书中有提到论文
typedef struct __Zem_t {
	int value;
	pthread_cond_t cond;
	pthread_mutex_t lock;
} Zem_t;

// only one thread can call this
void Zem_init(Zem_t *s, int value) {
	s->value = value;
	Cond_init(&s->cond);
	Mutex_init(&s->lock);
}

void Zem_wait(Zem_t *s) {
	Mutex_lock(&s->lock);
	while (s->value <= 0)
		Cond_wait(&s->cond, &s->lock);
	s->value--;
	Mutex_unlock(&s->lock);
}
void Zem_post(Zem_t *s) {
	Mutex_lock(&s->lock);
	s->value++;
	Cond_signal(&s->cond);
	Mutex_unlock(&s->lock);
}

HW:

5.reader-write-nostarve

void rwlock_acquire_readlock(rwlock_t *rw) {
    sem_wait(&rw->S);
    sem_post(&rw->S);
    sem_wait(&rw->lock);
    rw->readers++;
    if(rw->readers==1)
        sem_wait(&rw->writelock);
    sem_post(&rw->lock);
}

void rwlock_release_readlock(rwlock_t *rw) {
    sem_wait(&rw->lock);
    rw->readers--;
    if(rw->readers==0)
        sem_post(&rw->writelock);
    sem_post(&rw->lock);
}

void rwlock_acquire_writelock(rwlock_t *rw) {
    sem_wait(&rw->S);
    sem_wait(&rw->writelock);
}

void rwlock_release_writelock(rwlock_t *rw) {
    sem_post(&rw->S); //这行代码的位置有讲究,书上是放在这里,我觉得放在sem_post(&rw->writelock)后面或者sem_wait(&rw->writelock)前面好像都行
    sem_post(&rw->writelock);
}

6.no-starve-mutex

很难的问题,参考The Little Book of Semaphores 4.3节

weak semaphore和strong semaphore:strong semaphore能确保在一个wait线程之前唤醒的线程数量有界

no-starve-mutex的目的是基于weak semaphore实现no starving,具体实现非常巧妙,设两个room,轮流全部倒出,这样就不会出现单线程的loop。

t1、t2和mutex三个信号量,状态转移图如下:

进程状态转移

32.Common Concurrency Problems

CRUX: how to handle common concurrency bugs?
Non-Deadlock Bugs

1.atomicity-violation bugs

Thread 1::
if (thd->proc_info) {
	fputs(thd->proc_info, ...);
}

Thread 2::
thd->proc_info = NULL;

2.Order-Violation Bugs

用条件变量解决

Deadlock Bugs

一个死锁tutorial:https://deadlockempire.github.io/

CRUX: how to deal with deadlock?

为什么会有出现死锁?

  • large code bases, complex dependencies
  • encapsulation,底层细节,比如Java Vector class v1.AddAll(v2)需要multi-thread safe,获取锁的顺序随机

死锁条件

  • Mutual exclusion: Threads claim exclusive control of resources that they require (e.g., a thread grabs a lock)
  • Hold-and-wait: Threads hold resources allocated to them (e.g., locks that they have already acquired) while waiting for additional resources (e.g., locks that they wish to acquire)
  • No preemption: Resources (e.g., locks) cannot be forcibly removed from threads that are holding them
  • Circular wait: There exists a circular chain of threads such that each thread holds one or more resources (e.g., locks) that are being requested by the next thread in the chain.

Prevention:分别针对上面的条件

Circular wait

  • total ordering 固定锁的唤醒顺序
  • partial ordering: linux filemap.c
  • 小Tip:Enforce Lock Ordering by Lock Address
if (m1 > m2) { // grab in high-to-low address order
  pthread_mutex_lock(m1);
  pthread_mutex_lock(m2);
} else {
  pthread_mutex_lock(m2);
  pthread_mutex_lock(m1);
}// Code assumes that m1 != m2 (not the same lock)

Hold-and-wait

  • 用一个prevention锁包住所有锁的获取,意义不大

No Preemption

  • a deadlock-free, ordering-robust lock acquisition protocol
top:
	pthread_mutex_lock(L1);
	if (pthread_mutex_trylock(L2) != 0) {
    pthread_mutex_unlock(L1);
    goto top;
  }
  • 可能有livelock,但太巧了,可以用随机性处理
  • 存在encapsulation的问题,但这个解法至少对某些场合有效

Mutual Exclusion

  • lock-free系列方法,见第28节或笔记文件夹内threads-bugs.cpp文件
  • 感觉这类方法适用于分布式系统,backup、试错、no bounded loop

Deadlock Avoidance via Scheduling

  • 全局信息=>更优决策
  • Dijkstra’s Banker’s Algorithm [D64]
  • 应用不多:场景有限,比如嵌入式系统;限制了并行性
  • 设计理念: “Not everything worth doing is worth doing well”

Detect and Recover

A deadlock detector runs periodically, building a resource graph and checking it for cycles. 常应用于数据库

HW

./vector-*** -n 8 -d -l 100000 -t

vector-try-wait比vector-global-order速度略快

vector-avoid-hold-and-wait用global锁住local锁的获取,单线程有优势,但不支持并行操作

vector-nolock利用fetch-and-add,效率较低

33.Event-based Concurrency (Advanced)

CRUX: how to build concurrent servers without threads?

Event-based Concurrency

  • event handler独占时间,explicit control over scheduling
  • event-based servers中一定不要block!
while (1) {
	events = getEvents();
	for (e in events)
		processEvent(e);
}

an important API: select() (or poll())

int select(int nfds,fd_set *restrict readfds,fd_set *restrict writefds,fd_set *restrict errorfds,struct timeval *restrict timeout);
  • 这个api的意义是monitor各种fd是否“可用”(比如有新信息可读、有新空间可写)
  • timeout参数使用灵活,NULL表示允许无限block,0表示立刻返回,类似于waitpid的参数WNOHANG
  • pselect针对pthread做sigmask处理

A Problem: Blocking System Calls

=> no blocking calls are allowed => AIO

A Solution: Asynchronous I/O

思考:对于一个特殊的问题场景,需要厘清可用操作的边界,必要时可能引入新的概念,例如这里的asynchronous I/O、CSAPP p801的async-signal-safe functions

AIO control block: 利用aio_error配合signal机制interrupt,这一思想也用于I/O devices中

  • "Flash"这篇论文用hybrid思想,events are used to process network packets, and a thread pool is used to manage outstanding I/Os

Another Problem: State Management

manual stack management => use an old programming language construct known as a continuation, 用hash table这种数据结构存continuation信息

What Is Still Difficult With Events

  1. 不适用于多CPU
  2. 和systems activity配合不行,比如paging,是implicit blocking
  3. 不容易manage over time,改api的routine
  4. 这个系统的实现并不容易,hybrid

Appendix

Linux 系统文件

  • /etc/fstab file is a system configuration file that contains all available disks, disk partitions and their options
  • cat /proc/interrupts | grep "TLB shootdowns"
  • /proc/meminfo 可用内存
  • /proc/swaps 可swap内存
  • /proc/pid/fd/ 查找持有的fd,查文件泄漏
    • /proc/$tid 或 /prod/$pid/task/$tid 内核任务调度id
    • /proc/pid/status 线程数目
  • /proc/[pid]/uid_map 和 /proc/[pid]/gid_map
  • /proc/self/maps
  • /proc/sys/kernel/ns_last_pid
  • /proc/version 查看系统版本

TODO:


Linux多线程服务端编程 muduo

[toc]

muduo 是一个基于非阻塞 IO 和事件驱动的现代 C++ 网络库,原生支持 one loop per thread 这种 IO 模型。muduo 适合开发 Linux 下的面向业务的多线程服务端网络应用程序

Part I: C++ 多线程系统编程

chpt 1 线程安全的对象生命期管理

  • 基本问题是处理析构,用智能指针解决
  • 依据 [JCP],一个线程安全的 class 应当满足以下三个条件:
    • 多个线程同时访问时,其表现出正确的行为。
    • 无论操作系统如何调度这些线程,无论这些线程的执行顺序如何交织(interleaving)。
    • 调用端代码无须额外的同步或其他协调动作。
  • 对象创建:不要在构造函数泄漏this指针,最后一行也不行(因为接着可能执行派生类的代码)
    • 不要在构造函数中注册任何回调
  • 析构函数
    • 作为数据成员的 mutex 不能保护析构
    • 同时读写一个 class 的两个对象,有潜在的死锁可能
* 为了保证始终按相同的顺序加锁,我 们可以比较 mutex 对象的地址,始终先加锁地址较小的 mutex
  • 线程安全的 Observer 有多难
    • 对象的关系主要有三种:composition、aggregation、 association
    • 如果对象 x 注册了任何非静态成员函数回调,那么必然在某处持有了指向 x 的指针,这就暴露在了 race condition 之下
  • 原始指针有何不妥
    • 垃圾回收的原理,所有人都用不到的东西一定是垃圾
    • 解决思路:引入另外一层间接性(another layer of indirection)
  • 关于智能指针
    • shared_ptr/weak_ptr 的线程安全级别与 std::string 和 STL 容器一样($1.9)
* 并发读写shared_ptr(它本身,不是指它指向的对象),要加锁
* 销毁行为移出临界区:利用local_ptr和global_ptr做swap()
  • 原子操作,性能不错
* 二进制兼容性
* 讲解deleter传参的实现,模版相关:https://www.artima.com/articles/my-most-important-c-aha-momentsemeverem,泛型编程和面向对象编程的一次完美结合
  • 析构所在的线程:我们可以用一个单 独的线程来专门做析构,通过一个 BlockingQueue<shared_ptr<void> > 把对象的析 构都转移到那个专用线程,从而解放关键线程
  • 现成的 RAII handle
* 注意避免循环引用,通常的做法是 owner 持有指向 child 的 shared_ptr,child 持有指向 owner 的 weak_ptr
  • 对象池
    • 释放对象:weak_ptr
    • 解决内存泄漏:shared_ptr初始化传入delete函数
    • this指针线程安全问题:enable_shared_from_this
    • Factory生命周期延长:弱回调
* 通用的弱回调封装见 recipes/thread/WeakCallback.h,用到了 C++11 的 variadic template 和 rvalue reference
* 在事件通知中非常有用
  • 替代方案
    • 全局facade,对象访问都加锁,代价是性能,可以像Java的ConcurrentHashMap分buckets降低锁的代价
  • Observer 之谬
    • recipes/thread/SignalSlot.h
Note
  • 临界区在 Windows 上是 struct CRITICAL_SECTION,是可重入的;在 Linux 下是 pthread_mutex_t,默认是不可重入的
  • 在 Java 中,一个 reference 只要不为 null,它一定指向有效的对象
  • 如果这几种智能指针是对象 x 的数据成员,而它的模板参数 T 是个 incomplete 类型,那么 x 的析构函数不能是默认的或内联的,必须在 .cpp 文件里边 显式定义,否则会有编译错或运行错(原因见 §10.3.2)
    • 智能指针参考
  • * 过程范式、函数范式、对象范式
    • 对象范式的两个基本观念:
- 程序是由**对象**组成的;
- 对象之间互相**发送消息**,协作完成任务;
- 请注意,这两个观念与后来我们熟知的面向对象三要素“封装、继承、多态”根本不在一个层面上,倒是与再后来的“组件、接口”神合
  • C++的静态消息机制:不易开发windows这种动态消息场景;“面向类的设计”过于抽象
  • Java和.NET中分别对C++最大的问题——缺少对象级别的delegate机制做出了自己的回应
  • Realizing that C++’s “special” member functions may be declared private, 1988
  • Understanding the use of non-type template parameters in Barton’s and Nackman’s approach to dimensional analysis, 1995
* https://learningcppisfun.blogspot.com/2007/01/units-and-dimensions-with-c-templates.html
  • Understanding what problem Visitor addresses, 1996 or 1997
  • Understanding why remove doesn’t really remove anything, 1998
* remove不改变容器的元素个数,本质上是因为在STL中,container 和 algorithm (基于iterator) 在设计上分离的概念,比如当 algorithm 遇到了 array 类型,"couldn't change the size"
  • Understanding how deleters work in Boost’s shared_ptr, 2004.

chpt 2 线程同步精要

  • 并发编程有两种基本模型,一种是 message passing,另一种是 shared memory
  • 线程同步的四项原则,按重要性排列
    • 首要原则是尽量最低限度地共享对象,减少需要同步的场合。一个对象能不暴 露给别的线程就不要暴露;如果要暴露,优先考虑 immutable 对象;实在不行 才暴露可修改的对象,并用同步措施来充分保护它。
    • 其次是使用高级的并发编程构件,如 TaskQueue、Producer-Consumer Queue、 CountDownLatch 等等。
    • 最后不得已必须使用底层同步原语(primitives)时,只用非递归的互斥器和条件变量,慎用读写锁,不要用信号量。
    • 除了使用 atomic 整数之外,不自己编写 lock-free 代码,也不要用“内核级” 同步原语 。不凭空猜测“哪种做法性能会更好”,比如 spin lock vs. mutex。
  • 互斥器 mutex
  • Mutex
* 核心原则:利用RAII保证各项细则
* Scoped Locking:不手工调用 lock() 和 unlock() 函数,一切交给栈上的 Guard 对象的构造和析构函数负责
* 在每次构造 Guard 对象的时候,思考一路上(调用栈上)已经持有的锁,防止因加锁顺序不同而导致死锁(deadlock)
* 必要的时候可以考虑用 PTHREAD_MUTEX_ERRORCHECK 来排错
  • 只使用非递归的 mutex
* 概念:recursive or reentrant 
* recursive mutex 会隐藏代码问题:recipes/thread/test/NonRecursiveMutex_test.cc
  * post-mortem
* ```c++
  void postWithLockHold(const Foo& f) {
  	assert(mutex.isLockedByThisThread()); // muduo::MutexLock 提供了这个成员函数
  	// ... 
  }
  
    * 性能: 
  
      * Linux 的 Pthreads mutex 采用 futex(2) 实现,不必每次加锁、解锁都陷入系统调用,效率不错
        * [futex](https://akkadia.org/drepper/futex.pdf) TODO
  
      * Windows 的 [CRITICAL_SECTION](http://msdn.microsoft.com/en-us/library/windows/desktop/ms682530(v=vs.85).aspx) 也是类似的,不过它可以嵌入一小段 spin lock。在多 CPU 系统上,如果不能立刻拿到锁,它会先 spin 一小段时间,如果还不能拿到锁,才挂起当前线程
  
  * 死锁
  
    * `recipes/thread/test/SelfDeadLock.cc`
    * `recipes/thread/test/MutualDeadLock.cc`
  
  * false sharing and CPU cache TODO
  
    * http://www.aristeia.com/TalkNotes/ACCU2011_CPUCaches.pdf
    * http://www.akkadia.org/drepper/cpumemory.pdf
    * http://igoro.com/archive/gallery-of-processor-cache-effects/
    * http://simplygenius.net/Article/FalseSharing
  
* 条件变量

  * Java Object 内置的 wait()、notify()、 notifyAll() 是条件变量,以容易用错著称,一般建议用 java.util.concurrent 中的同步原语

  * ```c++
    // 经典应用 BlockingQueue
    muduo::MutexLock  mutex;
    muduo::Condition  cond(mutex);
    std::deque<int>   queue;
    int dequeue()
    {
      MutexLockGuard lock(mutex);
      while (queue.empty()) // 必须用循环;必须在判断之后再 wait()
      {
        cond.wait(); // 这一步会原子地 unlock mutex 并进入等待,不会与 enqueue 死锁
      // wait() 执行完毕时会自动重新加锁
      }
      assert(!queue.empty());
      int top = queue.front();
      queue.pop_front();
      return top;
    }
    
    void enqueue(int x)
    {
    	MutexLockGuard lock(mutex);
      queue.push_back(x);
    	cond.notify(); // 可以移出临界区之外
    }
  • broadcast should generally be used to indicate state change rather than resource availability
  • Mutex 和 Condition,就像与非门和 D 触发器构成了数字电路设计所需的全部基础元件,可以完成任何组合和同步时序逻辑电路设计一样
  • spurious wakeup: spurious wakeups can happen whenever there's a race and possibly even in the absence of a race or a signal
* 最简单的场景应该是有人给你的程序发了个Signal,当前正在执行的各种各样的系统调用都可能被中断。
* `cv.wait(lk, []{ return whether the event has occurred; });` 预防spurious wakeup
  * ```
    while (!pred()) {
        wait(lock);
    }

  * notify_one() 的细节,[讨论要不要放在锁里](https://en.cppreference.com/w/cpp/thread/condition_variable/notify_one)

    * "hurry up and wait" scenario
    * when precise scheduling of events is required

* 不要用读写锁和信号量

  * 读锁并不比 Mutex 高效
  * reader lock 可能允许提升(upgrade)为 writer lock,也可能不允许提升
  * 通常 reader lock 是可重入的,writer lock 是不可重入的。但是为了防止 writer 饥饿,writer lock 通常会阻塞后来的 reader lock,因此 **reader lock 在重入的时候可能死锁**。另外,在追求低延迟读取的场合也不适用读写锁,见 p. 55。
    * 注:java的ReentrantReadWriteLock实现,允许非公平锁,[读锁重入不会死锁](https://heapdump.cn/article/3957407),而Go会死锁
  * 性能问题怎么办?
    * 2.8 copy-on-write
    * read-copy-update
  * 信号量:哲学家就餐问题,“平权不如集权”
    * 如果要控制并发度,可以考虑用 muduo::ThreadPool
  * barrier原语:不如 CountDownLatch

* 封装 MutexLock、MutexLockGuard、Condition

* 线程安全的 Singleton 实现

  * C++:pthread_once、DCL(见本节Note)、Memory Barrier、Eager Initialization

* sleep(3) 不是同步原语

  * 生产代码中线程的等待可分为两种:一种是等待资源可用(要么等在 select/ poll/epoll_wait 上,要么等在条件变量上;一种是等着进入临界区(等在 mutex 上)以便读写共享数据。后一种等待通常极短,否则程序性能和伸缩性就会有问题
  * 在用户态做轮询(polling)是低效的

* 总结

  * 本文没有考虑 signal 对多线程编程的影响(§4.10),Unix 的 signal 在多线程下的行为比较复杂,一般要靠底层的网络库(如 Reactor)加以屏蔽,避免干扰上层应用程序的开发

* 借 shared_ptr 实现 copy-on-write

  * 用普通 mutex 替换读写锁的一个例子



##### Note

* [Real-world Concurrency](https://queue.acm.org/detail.cfm?id=1454462) 

  * history context
    * it was the introduction of the Burroughs B5000 in 1961 that proffered the idea that ultimately proved to be the way forward: disjoint CPUs concurrently executing different instruction streams but sharing a common memory
    * 1980: cache coherence protocols、prototyped parallel operating systems、parallel databases
    * 1990: symmetric multiprocessing,硬件+软件,uniprocessors to multiprocessors
      * microprocessor architects incorporated deeper (and more complicated) pipelines, caches, and prediction units
      * Many saw these two trends—the rise of concurrency and the futility of increasing clock rate—and came to the logical conclusion: instead of spending transistor budget on “faster” CPUs that weren’t actually yielding much in terms of performance gains (and had terrible costs in terms of power, heat, and area), why not take advantage of the rise of concurrent software and use transistors to effect multiple (simpler) cores per die?
      * That it was the success of concurrent software that contributed to the genesis of chip multiprocessing is an incredibly important historical point and bears reemphasis.
  * Concurrency is for Performance
    *  Just as no programmer felt a moral obligation to eliminate pipeline stalls on a superscalar microprocessor, no software engineer should feel responsible for using concurrency simply because the hardware supports it.
      * problems不一定应该/值得并行
    * to hide latency, for example, a disk I/O operation or a DNS lookup
      * 不一定值得
      * one can often achieve the same effect by employing nonblocking operations (e.g., asynchronous I/O) and an event loop (e.g., the poll()/select() calls found in Unix) in an otherwise sequential program.
    * to increase throughput need *not* consist exclusively (or even largely) of multithreaded code
      * e.g. typical MVC (model-view-controller) application
        * [为什么我不再推荐使用MVC框架?](https://www.toutiao.com/article/6763420080542843399/)
      * it is concurrency by architecture instead of by implementation.
  * Illuminating the Black Art
    * oral tradition in lieu of formal writing has left the domain shrouded in mystery
    * Know your cold paths from your hot paths.
    * Intuition is frequently wrong—be data intensive  ~ 压测
      * timeliness is more important than absolute accuracy: the absence of a perfect load simulation should not prevent you from simulating load altogether
      * Understanding scalability inhibitors on a production system requires the ability to safely dynamically instrument its synchronization primitives.
      * breaking up a lock is not the only way to reduce contention, and contention can be (and often is) more easily reduced by decreasing the hold time of the lock. This can be done by algorithmic improvements (many scalability improvements have been achieved by reducing execution under the lock from quadratic time to linear time!) or by finding activity that is needlessly protected by the lock
        * 经典例子:deallocate放到锁外
      * Be wary of readers/writer locks.
        * a readers/writer lock will use a single word of memory to store the number of readers
      * Consider per-CPU locking
        * if one were implementing a global counter that is frequently updated but infrequently read, one could implement a per-CPU counter protected by its own lock. Updates to the counter would update only the per-CPU copy, and in the uncommon case in which one wanted to read the counter, all per-CPU locks could be acquired and their corresponding values summed.
        * be sure to have a single order for acquiring all locks in the cold path
      * Know when to broadcast—and when to signal
        * Broadcast ~ *state change*, Signal ~ *resource availability*
        * *thundering herd*
      * Learn to debug postmortem
      * Second (and perhaps counterintuitively), one can achieve concurrency and composability by having no locks whatsoever. In this case, there must be no global subsystem state—subsystem state must be captured in per-instance state, and it must be up to consumers of the subsystem to assure that they do not access their instance in parallel. By leaving locking up to the client of the subsystem, the subsystem itself can be used concurrently by different subsystems and in different contexts
        * A concrete example of this is the AVL tree implementation used extensively in the Solaris kernel. As with any balanced binary tree, the implementation is sufficiently complex to merit componentization, but by not having any global state, the implementation may be used concurrently by disjoint subsystems—the only constraint is that manipulation of a single AVL tree instance must be serialized.
      * Don’t use a semaphore where a mutex would suffice.
        * unlike a semaphore, a mutex has a notion of *ownership*
        * First, there is no way of propagating the blocking thread’s scheduling priority to the thread that is in the critical section. This ability to propagate scheduling priority—*priority inheritance*—is critical in a realtime system, and in the absence of other protocols, semaphore-based systems will always be vulnerable to priority inversions.
      * Consider memory retiring to implement per-chain hash-table locks.
        * 见 【code-reading笔记】illumos部分
      * Be aware of false sharing.
        * This most frequently arises in practice when one attempts to defract contention with **an array of locks**
        * it can be expected to be even less of an issue on a multicore system (where caches are more likely to be shared among CPUs)
        * In this situation, array elements should be padded out to be a multiple of the coherence granularity.
      * Consider using nonblocking synchronization routines to monitor contention.
        * 见【code-reading笔记】illumos部分的per-cpu cache
      * When reacquiring locks, consider using generation counts to detect state change.
      * Use wait- and lock-free structures only if you absolutely must.
      * Prepare for the thrill of victory—and the agony of defeat.
  * The Concurrency Buffet
    * Those practitioners who are implementing a database or an operating system or a virtual machine will continue to need to sweat the details of writing multithreaded code
  
* [The "Double-Checked Locking is Broken" Declaration](http://www.cs.umd.edu/~pugh/java/memoryModel/DoubleCheckedLocking.html)

  * ```java
    / Broken multithreaded version
    // "Double-Checked Locking" idiom
    class Foo { 
      private Helper helper = null;
      public Helper getHelper() {
        if (helper == null) 
          synchronized(this) {
            if (helper == null) 
              helper = new Helper();
          }    
        return helper;
      }
      // other functions and members...
    }
  • Unfortunately, that code just does not work in the presence of either optimizing compilers or shared memory multiprocessors. There is no way to make it work without requiring each thread that accesses the helper object to perform synchronization.
* obj产生和构造函数在inline后可能乱序,getHelper()拿到未构造完成的对象
  • A fix that doesn't work
* The rule for a monitorexit (i.e., releasing synchronization) is that actions before the monitorexit must be performed before the monitor is released. However, there is no rule which says that actions after the monitorexit may not be done before the monitor is released. 
  • 必须加 memory_barrier,但还不够
* The problem is that on some systems, the thread which sees a non-null value for the `helper` field also needs to perform memory barriers.
* 本质上是需要 cache coherence instruction
  • Making it work for static singletons
  • It will work for 32-bit primitive values
* it does not work for long's or double's, since unsynchronized reads/writes of 64-bit primitives are not guaranteed to be atomic.
* ```c++
  // Lazy initialization 32-bit primitives
  // Thread-safe if computeHashCode is idempotent
  class Foo { 
    private int cachedHashCode = 0;
    public int hashCode() {
      int h = cachedHashCode;
      if (h == 0) {
        h = computeHashCode();
        cachedHashCode = h;
      }
      return h;
    }
    // other functions and members...
  }

  * ```c++
    // C++ implementation with explicit memory barriers
    // Should work on any platform, including DEC Alphas
    // From "Patterns for Concurrent and Distributed Objects",
    // by Doug Schmidt
    template <class TYPE, class LOCK> TYPE *
    Singleton<TYPE, LOCK>::instance (void) {
        // First check
        TYPE* tmp = instance_;
        // Insert the CPU-specific memory barrier instruction
        // to synchronize the cache lines on multi-processor.
        asm ("memoryBarrier");
        if (tmp == 0) {
            // Ensure serialization (guard constructor acquires lock_).
            Guard<LOCK> guard (lock_);
            // Double check.
            tmp = instance_;
            if (tmp == 0) {
                    tmp = new TYPE;
                    // Insert the CPU-specific memory barrier instruction
                    // to synchronize the cache lines on multi-processor.
                    asm ("memoryBarrier");
                    instance_ = tmp;
            }
        }
        return tmp;
    }
  • Fixing Double-Checked Locking using Thread Local Storage
  • Under the new Java Memory Model
* Fixing Double-Checked Locking using Volatile
* Double-Checked Locking Immutable Objects
  • durations的角度
* For short lock durations, up to say 10%, the system achieved very high parallelism. Not perfect parallelism, but close. Locks are fast!
* once the lock duration passes 90%, there’s no point using multiple threads anymore
  • lock frequency的角度:
* As my [next post](http://preshing.com/20111124/always-use-a-lightweight-mutex) shows, a pair of lock/unlock operations on a Windows Critical Section takes about **23.5 ns** on the CPU used in these tests
* us 级别时,综合性能表现较好
  • the lock around the memory allocator in a game engine will often achieve excellent performance.
  • 本文没有结合 CPU usage 做分析,最好能定量 CPU 损失以及连带的吞吐影响

chpt 3 多线程服务器的适用场合与常用编程模型

  • 进程与线程
  • 《Erlang 程序设计》[ERL] 把“进程” 比喻为“人”,我觉得十分精当,为我们提供了一个思考的框架。
* 每个人有自己的记忆(memory),人与人通过谈话(消息传递)来交流,谈话既可以是面谈(同一台服务器),也可以在电话里谈(不同的服务器,有网络通信)。
* 面谈和电话谈的区别在于,面谈可以立即知道对方是否死了(crash, SIGCHLD),而电话谈只能通过周期性的心跳来判断对方是否还活着。
  • 线程的特点是共享地址空间,从而可以高效地共享数据。一台机器上的多个进程 能高效地共享代码段(操作系统可以映射为同样的物理内存),但不能共享数据。如果多个进程大量共享内存,等于是把多进程程序当成多线程来写,掩耳盗铃。
  • 单线程服务器的常用编程模型
  • Reactor模式:“non-blocking IO + IO multiplexing”
* lighttpd,单线程服务器。(Nginx与之类似,每个工作进程有一个eventloop。) 
* libevent,libev。ACE,Poco C++ libraries。
* Java NIO,包括 Apache Mina 和 Netty。POE(Perl)。Twisted (Python)。
* 优点:不仅可以用于读写 socket, 连接的建立(connect(2)/accept(2))甚至 DNS 解析 4 都可以用非阻塞方式进行,以 提高并发度和吞吐量(throughput),对于 IO 密集的应用是个不错的选择
* 缺点:它要求事件回调函数必须是非阻塞 的。对于涉及网络 IO 的请求响应式协议,它容易割裂业务逻辑,使其散布于多个回调函数之中,相对不容易理解和维护
* ```c++
  while (!done) {
    int timeout_ms = max(1000, getNextTimedCallback());
    int retval = ::poll(fds, nfds, timeout_ms);
    if (retval < 0) {
      处理错误,回调用户的 error handler
    } else {
      处理到期的 timers,回调用户的 timer handler
        if (retval > 0) {
          处理 IO 事件,回调用户的 IO event handler }
    	}
  	}
  }

  * Proactor模式:

    * Boost.Asio 和 Windows I/O Completion Ports
    * 和Reactor模式的区别在于由内核完成IO操作(同步IO可以模拟异步IO),用户业务逻辑无阻塞

* 多线程服务器的常用编程模型

  * non-blocking IO + one loop per thread:处理IO和定时器
    * Event loop 代表了线程的主循环,需要让哪个线程干活,就把 timer 或 IO channel (如 TCP 连接)注册到哪个线程的 loop 里即可。对实时性有要求的 connection 可以单独用一个线程;数据量大的 connection 可以独占一个线程,并把数据处理任务分摊到另几个计算线程中(用线程池);其他次要的辅助性 connections 可以共享一个线程。
    * 线程安全很重要
  * 线程池:处理计算
    * 任务队列 或 生产者消费者数据队列
      * `concurrent_queue<T>`
    * “阻抗匹配”(p80)

* 进程间通信只用 TCP

  * pipe 也有一个经典应用场景,那就是写 Reactor/event loop 时用来[异步唤醒 select (或等价的 poll/epoll_wait)调用](https://www.zhihu.com/question/39752285/answer/82906915)
    * 在 Linux 下,可以用 eventfd(2) 代替,效率更高
    * elf pipe trick,说白了也很简单,windows下select只能针对socket套接字,不能针对管道,一般用构造两个互相链接于[localhost](https://www.zhihu.com/search?q=localhost&search_source=Entity&hybrid_search_source=Entity&hybrid_search_extra={"sourceType"%3A"answer"%2C"sourceId"%3A82906915})的socket来模拟之。不过win下select最多支持同时wait 64个套接字,你摸拟的[pipe](https://www.zhihu.com/search?q=pipe&search_source=Entity&hybrid_search_source=Entity&hybrid_search_extra={"sourceType"%3A"answer"%2C"sourceId"%3A82906915})占掉一个,就只剩下63个可用了。所以java的nio里[selector](https://www.zhihu.com/search?q=selector&search_source=Entity&hybrid_search_source=Entity&hybrid_search_extra={"sourceType"%3A"answer"%2C"sourceId"%3A82906915})在windows下最多支持62个套接字就是被self pipe trick占掉了两个,一个用于其它线程调用notify唤醒,另一个留作jre内部保留,就是这个原因。
  * Note
    * 用 socket_pair(2) 做双向通信
    * 消息格式:推荐Protobuf
    * TCP 的 local 吞吐量不低
    * 另外,除了点对点的通信之外,应用级的广播协议也是非常有用的,可以方便地构建可观可控的分布式系统,见 §7.11
  * 分布式系统中使用 TCP 长连接通信
    * 容易定位分布式系统中的服务之间的依赖关系
    * 通过接收和发送队列的长度也较容易定位网络或程序故障

* 多线程服务器的适用场合

  * 两种方式:宝贵的原生线程(reactor模式,pthread_create),廉价的“线程”(阻塞io,语言runtime调度)

    * Pthreads 是 NPTL(Native POSIX Thread Library) 的,每个线程由 clone(2) 产生,对应一个内核的 task_struct

  * 使用速率为 50MB/s 的数据压缩库、在进程创建销毁 的开销是 800μs、线程创建销毁的开销是 50μs 的前提下,考虑如何执行压缩任务:

    - 如果要偶尔压缩 1GB 的文本文件,预计运行时间是 20s,那么起一个进程去做是合理的,因为进程启动和销毁的开销远远小于实际任务的耗时。
    - 如果要经常压缩 500kB 的文本数据,预计运行时间是 10ms,那么每次都起进程似乎有点浪费了,可以每次单独起一个线程去做。
    - 如果要频繁压缩 10kB 的文本数据,预计运行时间是 200μs,那么每次起线程似乎也很浪费,不如直接在当前线程搞定。也可以用一个线程池,每次把压缩任务交给线程池,避免阻塞当前线程(特别要避免阻塞 IO 线程)。

  * 必须用单线程的场合

    * 程序可能会 fork(2)
      * 立刻执行 exec(),变身为另一个程序。例如 shell 和 inetd; 又比如 lighttpd fork() 出子进程,然后运行 fastcgi 程序。或者集群中运行在计算节点上的负责启动 job 的守护进程(即所谓的“看门狗进程”)。
      * 不调用 exec(),继续运行当前程序。要么通过共享的文件描述符与父进程通信,协同完成任务;要么接过父进程传来的文件描述符,独立完成工作,例如 20 世纪 80 年代的 Web 服务器 NCSA httpd。
    * 限制程序的 CPU 占用率
      * 因此对于一些辅助性的程序,如果它必须和主要服务进程运行在同一台机器的话 (比如它要监控其他服务进程的状态),那么做成单线程的能避免过分抢夺系统的计算资源。比方说如果要把生产服务器上的日志文件压缩后备份到 NFS 上,那么应该使用普通单线程压缩工具(gzip/bzip2)。它们对系统造成的影响较小,在 8 核服务器上最多占满 1 个 core。

  * 单线程程序的优缺点

    * Event loop 有一个明显的缺点,它是非抢占的(non-preemptive)。这个缺点可以用多线程来克服
    * [IOCP , kqueue , epoll ... 有多重要?](https://blog.codingnow.com/2006/04/iocp_kqueue_epoll.html)
      * 这篇blog讲逻辑服务器前面加一个gateway;gateway定时发数据利于逻辑服务器调试

  * 适用多线程程序的场景

    * 提高响应速度,让 IO 和“计算”相互重叠,降低latency。虽然多线程不能提高绝对性能,但能提高平均响应性能。
    * 多线程间有需要修改的共享数据,提供非均质的服务(对于高优任务防止优先级反转)
    * latency 和 throughput 同样重要,利用异步操作
    * 性能可预测,多线程能有效地划分责任与功能

  * 例子:

    * master-slave,master多线程
      * 4 个用于和 slaves 通信的 IO 线程。
      * 1 个 logging 线程。
      * 1 个数据库 IO 线程。
      * 2 个和 clients 通信的 IO 线程。
      * 1 个主线程,用于做些背景工作,比如 job 调度。
      * 1 个 pushing 线程,用于主动广播机群的状态。
    * TCP聊天服务器:转发连接,更多功能
      * 见 §6.6 的方案 9,以及 p. 260 的实现

  * “多线程服务器的适用场合”例释与答疑

    * Linux 能同时启动多少个线程?

      * 对于 32-bit Linux,一个进程的地址空间是 4GiB,其中用户态能访问 3GiB 左右, 而一个线程的默认栈(stack)大小是 10MB,心算可知,一个进程大约最多能同时启动 300 个线程

    * 多线程能提高并发度吗?

      * thread per connection 不适合高并发场合,其 scalability 不佳。one loop per thread 的并发度足够大,且与 CPU 数目成正比。

    * 多线程能提高吞吐量吗?

      * 对于计算密集型服务,不能。
      * 根据 Amdahl’s law,即便算法的并行度高达 95%,8 核的加速比也只有 6,计算 时间为 0.133s,这样会造成吞吐量下降
      * 线程池也不是万能的,如果响应一次请求需要做比较多的计算(比如计算的时间占整个 response time 的 1/5 强),那么用线程池是合理的,能简化编程。如果在一次请求响应中,主要时间是在等待 IO,那么为了进一步提高吞吐量,往往要用其他编程模型,比如 Proactor,见问题 8

    * 多线程能降低响应时间吗?

      * 多线程处理输入:并行化IO部分,降低平均延时(减少串行时某些任务的IO等待)
      * 多线程分担负载

    * 多线程程序如何让 IO 和“计算”相互重叠,降低 latency?

      * 所有的网络写操作都可以这么异步地做,不过这也有一个缺点,那就是每次 asyncWrite() 都要在线程间传递数据。其实如果 TCP 缓冲区是空的,我们就可以在本线程写完,不用劳烦专门的 IO 线程。Netty 就使用了这个办法来进一步降低延迟。

    * 第三方库不一定能很好地适应并融入这个 event loop framework

      * 但是检测串口上的某些控制信号(例如 DCD)只能用轮询(ioctl(fd, TIOCMGET, &flags))或阻塞等待(ioctl(fd, TIOCMIWAIT, TIOCM_CAR));要想融入 event loop,需要单独起一个线程来查询串口信 号翻转,再转换为文件描述符的读写事件(可以通过 pipe(2))
      * libmemcached 只支持同步操作

    * 什么是线程池大小的阻抗匹配原则?

      * T = C/P :密集计算所占的时间比重为 P (0 < P ≤ 1),而 系统一共有 C 个 CPU,为了让这 C 个 CPU 跑满而又不过载,线程池大小的经验公式

    * 除了你推荐的 Reactor + thread poll,还有别的 non-trivial 多线程编程模型吗?

      * Proactor 模式依赖操作系统或库来高效地调度这些子任务,每个子任务都不会阻

        塞,因此能用比较少的线程达到很高的 IO 并发度。

      * Proactor 能提高吞吐,但不能降低延迟,所以我没有深入研究。另外,在没有语 言直接支持的情况下 26,Proactor 模式让代码非常破碎,在 C++ 中使用 Proactor 是 很痛苦的。因此最好在“线程”很廉价的语言中使用这种方式,这时 runtime 往往会 屏蔽细节,程序用单线程阻塞 IO 的方式来处理 TCP 连接

    * 模式 2 和模式 3a 该如何取舍?

      * 可以根据工作集(work set)的大小来取舍。 工作集是指服务程序响应一次请求所访问的内存大小
      * memcached 这个内存消耗大户用多线程服务端就比在同一台机器上运行多个 memcached instance 要好。(但是如果你在 16GiB 内存的机器上运行 32-bit memcached,那么此时多 instance 是必需的。)
        * 地址空间4GiB,堆栈大小受限;单进程hack成多份地址空间No,多进程指定共享内存Yes

#### chpt 4 C++ 多线程系统编程精要

* 基本线程原语的选用

  * Thread + MutexLock + Condition
  * pthread_once,封装为 muduo::Singleton。其实不如直接用全局变量。
  * pthread_key*,封装为 muduo::ThreadLocal。可以考虑用 __thread 替换之。
  * 不建议使用:
    * pthread_rwlock,读写锁通常应慎用。muduo 没有封装读写锁,这是有意的。
    * sem\_*,避免用信号量(semaphore)。它的功能与条件变量重合,但容易用错。
    * pthread\_{cancel, kill}。程序中出现了它们,则通常意味着设计出了问题。

* C/C++ 系统库的线程安全性

  * 线程的出现立刻给系统函数库带来了冲击,破坏了 20 年来一贯的编程传统和假定。例如:

    * errno 不再是一个全局变量,因为每个线程可能会执行不同的系统库函数。

      * ```c++
        extern int *__errno_location(void);
        // return a lvalue
        #define errno (*__errno_location())
* 有些“纯函数”不受影响,例如 memset/strcpy/snprintf 等等。
* 有些影响全局状态或者有副作用的函数可以通过加锁来实现线程安全,例如malloc/free、printf、fread/fseek 等等。
  * printf线程安全、cout不线程安全
  * 非线程安全的性能更好的版本:fread_unlocked、fwrite_unlocked 等等,见 `man unlocked_stdio`
  * 例如 fseek() 和 fread() 都是安全的,但是对某个文件“先 seek 再 read”这两步操作中间有可能 会被打断,其他线程有可能趁机修改了文件的当前位置,让程序逻辑无法正确执行。 在这种情况下,我们可以用 flockfile(FILE*) 和 funlockfile(FILE*) 函数来显式地 加锁。并且由于 FILE* 的锁是可重入的,加锁之后再调用 fread() 不会造成死锁。
  * 如果程序直接使用 lseek(2) 和 read(2) 这两个系统调用来随机读取文件,也存 在“先 seek 再 read”这种 race condition,但是似乎我们无法高效地对系统调用加 锁。解决办法是改用 pread(2) 系统调用,它不会改变文件的当前位置。
* POSIX标准列出 [非线程安全函数的黑名单](https://pubs.opengroup.org/onlinepubs/9699919799/functions/V2_chap02.html#tag_15_09)
  * 有些返回或使用静态空间的函数不可能做到线程安全,因此要提供另外的版本,例如 asctime_r/ctime_r/gmtime_r、stderror_r、strtok_r 等等。
* 传统的 fork() 并发模型不再适用于多线程程序(§4.9)。
* 我们不必担心系统调用的线程安全性,因为系统调用对于用户态程序来说是原子的。但是要注意系统调用对于内核状态的改变可能影响其他线程,这个话题留到 §4.6 再细说。
  • 编写线程安全程序的一个难点在于线程安全是不可组合的(composable)
* e.g. `tzset()` 是全局的,会影响其它线程的时区状态 ---> `muduo::TimeZone`
* 一个基本思路是尽量把 class 设计成 immutable 的,这样用起来就不必为线程安全操心了
* C++ 标准库中的绝大多数泛型算法是线程安全的,因为这些都是无状态纯函数。只要输入区间是线程安全的,那么泛型函数就是线程安全的
  * `std::random_shuffle()` 可能是个例外,它用到了随机数发生器
  * 随机数发生不是线程安全的,因为`time(NULL)`随机种子可能一样,导致产生的随机数一样。因此best practice是`static thread_local std::mt19937 rng(std::random_device{}());`
* C++ 的 iostream 不是线程安全的,因为流式输出是多个operator <<函数调用
  * 如果用`printf`等于加了全局锁
  • Linux 上的线程标识
  • pthread_t 并不适合用作程序中对线程的标识符。
  • 在 Linux 上,我建议使用 gettid(2) 系统调用的返回值作为线程 id
* 在现代 Linux 中,它直接表示内核的任务调度 id,因此在 /proc 文件系统中可以轻易找到对应项:`/proc/$tid` 或 `/prod/$pid/task/$tid`
* 任何时刻都是全局唯一的,并且由于 Linux 分配新 pid 采用递增轮回办法,短时间内启动的多个线程也会具有不同的线程 id。
* 0 是非法值,因为操作系统第一个进程 init 的 pid 是 1
  • 线程的创建与销毁的守则
    • 几条创建的原则
* 程序库不应该在未提前告知的情况下创建自己的“背景线程”。
  * 一旦程序中有不止一个线程,就很难安全地 fork() 了 (§4.9)
* 尽量用相同的方式创建线程,例如 muduo::Thread。
  * bookkeeping,线程数目可以从 `/proc/pid/status` 拿到
* 在进入 main() 函数之前不应该启动线程。
  * C++ 保证在进入 main() 之前完成全局对象的构造
    * "全局对象"也包括 namespace 级全局对象、文件级静态对象、class 的静态对象,但不包 括函数内的静态对象
  * 如果一个库需要创建线程,那么应该进入 main() 函数之后再调用库的初始化函数去做
* 程序中线程的创建最好能在初始化阶段全部完成。
  • 线程的销毁有几种方式
* 自然死亡。从线程主函数返回,线程正常退出。
* 非正常死亡。从线程主函数抛出异常或线程触发 segfault 信号等非法操作
* 自杀。在线程中调用 pthread_exit() 来立刻退出线程。
* 他杀。其他线程调用 pthread_cancel() 来强制终止某个线程。
  * pthread_kill() 是往线程发信号,留到 §4.10 再讨论
  * 不要他杀!
  * 如果确实需要强行终止一个耗时很长的计算任务,而又不想在计算期间周期性 地检查某个全局退出标志,那么可以考虑把那一部分代码 fork() 为新的进程,这样杀(kill(2))一个进程比杀本进程内的线程要安全得多。当然,fork() 的新进程与 本进程的通信方式也要慎重选取,最好用文件描述符(pipe(2)/socketpair(2)/TCP socket)来收发数据,而不要用共享内存和跨进程的互斥器等 IPC,因为这样仍然有死锁的可能。
  • pthread_cancel 与 C++
* [Cancellation and C++ Exceptions](https://udrepper.livejournal.com/21541.html)
  * catch-all cases must rethrow
  * `#define CATCHALL catch (abi::__forced_unwind&) { throw; } catch (...)`
  • exit(3) 在 C++ 中不是线程安全的
* exit(3) 函数在 C++ 中的作用除了终止进程,还会析构全局对象和已经构造完的函数静态对象。这可能导致:死锁、其它线程调用已经被析构的全局对象
* 如果确实需要主动结束 线程,则可以考虑用 _exit(2) 系统调用。它不会试图析构全局对象,但是也不会执 行其他任何清理工作,比如 flush 标准输出。
  • 善用 __thread 关键字
    • 比 pthread_key_t 快很多
    • __thread 使用规则 27:只能用于修饰 POD 类型,不能修饰 class 类型,因为无法 自动调用构造函数和析构函数。__thread 可以用于修饰全局变量、函数内的静态变 量,但是不能用于修饰函数的局部变量或者 class 的普通成员变量。另外,__thread 变量的初始化只能用编译期常量。
    • 注意与C++ threadlocal 关键字比较
    • 书里举了一些应用的例子(p97)
  • 多线程与 IO
    • 网络IO
* 多个线程同时操作同一个 socket 文件描述符确实很麻烦,chenshuo认为是得不偿失的
* 各种read/write/connect/clost/listen的情况,太复杂了,而且read/write返回字节数也是不定的
  • 磁盘IO
* 要避免 lseek(2)/ read(2) 的 race condition(§4.2)
* 每块磁盘都有一个操作队列,多个线程的读写请求 到了内核是排队执行的。只有在内核缓存了大部分数据的情况下,多线程读这些热数据才可能比单线程快
  • epoll
* epoll 也遵循相同的原则。Linux 文档并没有说明:当一个线程正阻塞在 epoll_ wait() 上时,另一个线程往此 epoll fd 添加一个新的监视 fd 会发生什么。
* `muduo::EventLoop::wakeup()`
  • 为了简单起见,我认为多线程程序应该遵循的原则是:每个文件描述符只由一个线程操作,从而轻松解决消息收发的顺序性问题,也避免了关闭文件描述符的各种 race condition
  • 这条规则有两个例外:
* 对于磁盘文件,在必要的时候多个线程可以同时调用 pread(2)/pwrite(2) 来读写同一个文件;
* 对于 UDP,由于协议本身保证消息的原子性,在适当的条件下(比如消息之间彼此独立)可以多个线程同时读写同一个 UDP 文件描述符。--->  [相关讨论](https://www.zhihu.com/question/39185963)
  • 用 RAII 包装文件描述符
    • POSIX 标准要求每次新打开文件(含 socket)的时候必须使用当前最小可用的文件描述符号码。在多线程程序中,这样很容易串话
    • 在 C++ 里解决这个问题的办法很简单:RAII
    • 引申问题:为什么服务端程序不应该关闭标准输出(fd=1)和标准错误(fd=2)?
* 因为有些第三方库在特殊紧急情况下会往 stdout 或 stderr 打印出错信息,如果我们 的程序关闭了标准输出(fd=1)和标准错误(fd=2),这两个文件描述符有可能被网络连接占用,结果造成对方收到莫名其妙的数据。正确的做法是把 stdout 或 stderr 重定向到磁盘文件(最好不要是 /dev/null),这样我们不至于丢失关键的诊断信息。 当然,这应该由启动服务程序的看门狗进程完成,对服务程序本身是透明的
  • muduo 使用 shared_ptr 来管理 TcpConnection 的生命期。这是唯一一个采用引用计数方式管理生命期的对象。如果不用 shared_ptr,我想不出其他安全且高效的办法来管理多线程网络服务端程序中的并发连接
* 父进程的内存锁,mlock(2)、mlockall(2)。
* 父进程的文件锁,fcntl(2)。
* 父进程的某些定时器,setitimer(2)、alarm(2)、timer_create(2) 等等。
* 其他,见 man 2 fork。
  • 多线程与 fork()
    • 多线程与 fork() 的协作性很差。这是 POSIX 系列操作系统的历史包袱,因为以前长期是单线程的设计
* 无法forkall
  • 在 fork() 之后,子进程就相当于处于 signal handler 之中,你不能调用线程安全的函数(除 非它是可重入的),而只能调用异步信号安全(async-signal-safe)的函数
  • 唯一安全的做法是在 fork() 之后立即调用 exec() 执行另一个程序, 彻底隔断子进程与父进程的联系。
* 不得不说,同样是创建进程,Windows 的 CreateProcess() 函数的顾虑要少得多,因为它创建的进程跟当前进程关联较少。
  • 多线程与 signal
    • 单线程时代:由于 signal 打断了正在运行的 thread of control,在 signal handler 中只能调用 async-signal-safe 的函数,即 所谓的“可重入(reentrant)”函数,就好比在 DOS 时代编写中断处理例程(ISR)一样。不是每个线程安全的函数都是可重入的。
* the **First-Level Interrupt Handler** (**FLIH**) and the **Second-Level Interrupt Handlers** (**SLIH**). FLIHs are also known as *hard interrupt handlers* or *fast interrupt handlers*, and SLIHs are also known as *slow/soft interrupt handlers*, or [Deferred Procedure Calls](https://en.wikipedia.org/wiki/Deferred_Procedure_Call) in Windows.
  * SLIH use kernel threads
* 如果 signal handler 中需要修改全局数据,那么被修改的变量必须是 `sig_atomic_t`
  • 多线程时代
* 发送给某一线程(SIGSEGV),发送给进程中的任一线程(SIGTERM)
* 在多线程程序中,使用 signal 的第一原则是**不要使用 signal**
* 不主动处理各种异常信号(SIGTERM、SIGINT 等等),只用默认语义:结束进程。 有一个例外:SIGPIPE,服务器程序通常的做法是忽略此信号 40,否则如果对方 断开连接,而本机继续 write 的话,会导致程序意外终止
* 在没有别的替代方法的情况下(比方说需要处理 SIGCHLD 信号),把异步信号转换为同步的文件描述符事件。现代 Linux 的做法是采用 signalfd(2) 把信号直接转换为文件描述符事件,从而从根本上避免使用 signal handler
  * 例子见 http://github.com/chenshuo/muduo-protorpc 中 Zurg slave 示例的 [ChildManager class](https://github.com/chenshuo/muduo-protorpc/blob/cpp11/examples/zurg/slave/ChildManager.cc)
  • Linux 新增系统调用的启示
    • 大致从 Linux 内核 2.6.27 起,凡是会创建文件描述符的 syscall 一般都增加了额外的 flags 参数,可以直接指定 O_NONBLOCK 和 FD_CLOEXEC
* accept4 - 2.6.28, eventfd2 - 2.6.27, inotify_init1 - 2.6.27, pipe2 - 2.6.27, signalfd4 - 2.6.27, timerfd_create - 2.6.25
  • 另外,以下新系统调用可以在创建文件描述符时开启 FD_CLOEXEC 选项:
* 以前需要 `fcntl(fd, F_SETFD, FD_CLOEXEC);`
* open, dup3 - 2.6.27, epoll_create1 - 2.6.27, socket - 2.6.27
* [Secure File Descriptor Handling](https://udrepper.livejournal.com/20407.html)
  * 例子:web fork出plugins执行exec,不希望主进程已有的私密文件泄露给第三方
  * fork()完立刻set flag并不安全,因为fork()是signal-safe的 (i.e., it can be called from a signal handler).
  • Note
    • 在多 CPU 机器上,假设主板上两个物理 CPU 的距离为 15cm,CPU 主频是 2.4GHz,电信号在电路中 的传播速度按 2 × 108m/s 估算,那么在 1 个时钟周期(0.42ns)之内,电信号不能从一个 CPU 到达另一个 CPU。因此对于每个 CPU 自己这个观察者来说,它看到的事件发生的顺序没有全局一致性
    • 在现代 Linux glibc 中,fork(3) 不是直接使用 fork(2) 系统调用,而是使用 clone(2) syscall

chpt 5 高效的多线程日志

  • logging
  • 诊断日志(diagnostic log) 即 log4j、logback、slf4j、glog、g2log、log4cxx、 log4cpp、log4cplus、Pantheios、ezlogger 等常用日志库提供的日志功能
  • 交易日志(transaction log) 即数据库的 write-ahead log、文件系统的 journaling 等,用于记录状态变更,通过回放日志可以逐步恢复每一次修改之后的状态
  • 前端风格
* C/Java 的 printf(fmt, ...) 风格
  * printf(fmt, ...) 风格在 C++ 中也可以做到类型安全,但是在 C++11 引入 variadic template 之前很费 劲。因为 C++ 不允许把 non-POD 对象通过可变参数(...)传入函数。Pantheios 日志库用的是重载函数模板的办法(http://www.pantheios.org)
* C++ 的 stream << 风格
  * 用起来更自然,不必费心保持格式字符串与参数类型的一致性,可以随用随写,而且是类型安全的
  * stream 风格的另一个好处是当输出的日志级别高于语句的日志级别时,打印日志是个空操作,运行时开销接近零
  • 功能需求
  • 调整日志的输出级别不需要重新编译,也不需要重启进程,只要调用muduo::Logger::setLogLevel() 就能即时生效
  • 对于分布式系统中的服务进程而言,日志的目的地(destination)只有一个: 本地文件。往网络写日志消息是不靠谱的,因为诊断日志的功能之一正是诊断网络故障,比如连接断开(网卡或交换机故障)、网络暂时不通(若干秒之内没有收到心跳 消息)、网络拥塞(消息延迟明显加大)等等
  • 日志rolling
* 条件通常有两个:文件大小(例如每写满 1GB 就换下一个文件)和时间(例如每天零点新建一个日志文件,不论前一个文件有没有写满)
  • 日志文件压缩与归档 (archive)不是日志库应有的功能,而应该交给专门的脚本去做,这样 C++ 和 Java 的服务程序可以共享这一基础设施
  • 磁盘空间监控也不是日志库的必备功能:磁盘报警人工干预
  • 往文件写日志的一个常见问题是,万一程序崩溃,那么最后若干条日志往往就丢失了,因为日志库不能每条消息都 flush 硬盘,更不能每条日志都 open/close 文件,这样性能开销太大。muduo 日志库用两个办法来应对这一点,其一是定期(默认 3 秒)将缓冲区内的日志消息 flush 到硬盘;其二是每条内存中的日志消息都带有 cookie(或者叫哨兵值/sentry),其值为某个函数的地址,这样通过在 core dump 文件中查找 cookie 就能找到尚未来得及写入磁盘的消息。
* 可以用 gdb 的 find 命令。用 strings(1) 命令也能从 core 文件里找到不少有用的信息
  • 日志每行带上线程号
  • 时间戳精确到微秒。每条消息都通过 gettimeofday(2) 获得当前时间,这么做不会有什么性能损失。因为在 x86-64 Linux 上,gettimeofday(2) 不是系统调用,不会陷入内核 (可用 strace(1) 验证 muduo/base/tests/Timestamp_unittest.cc)
* [On vsyscalls and the vDSO](https://lwn.net/Articles/446528/):the kernel allows the page containing the current time to be mapped read-only into user space; that page also contains a fast `gettimeofday()` implementation
* ```shell
  $ cat /proc/self/maps
  ...
  7fffcbcb7000-7fffcbcb8000 r-xp 00000000 00:00 0            [vdso]
  ffffffffff600000-ffffffffff601000 r-xp 00000000 00:00 0    [vsyscall]

    * 内核代码和数据总是可寻址,随时准备处理中断和系统调用。与此相反,用户模式地址空间的映射随进程切换的发生而不断变化,Address-space layout randomization is a form of defense against security holes。

    * 但vsyscall的地址是固定的,可能被攻击 ---> The result is a kernel system call emulating a virtual system call which was put there to avoid the kernel system call in the first place ---> `CONFIG_UNSAFE_VSYSCALLS`

    * One useful point from that discussion is that the static vsyscall page is not, in fact, a security vulnerability; it's simply a resource which can make it easier for an attacker to exploit a vulnerability elsewhere in the system.

  * 始终使用 GMT 时区

  * 应该避免在日志格式(特别是消息 id )中出现正则表达式的元字符(meta character),例如 '[' 和 ']' 等等,这样在用 less(1) 查看日志文件的时候查找字符串更加便捷。

    * 对于 Base64 编码的消息 id,可以将其中的 '+' 替换为 '-',见 RFC 4648 第 5 节

* 性能需求

  * 1GB/min, 上万qps
  * 磁盘带宽约是 110MB/s,日志库应该能瞬时写满这个带宽
  * 假如每条日志消息的平均长度是 110 字节,这意味着 1 秒要写 100 万条日志
  * 性能优化见【code-reading笔记】

* 多线程异步日志

  * 在多线程服务程序中,异步日志(叫“非阻塞日志”似乎更准确)是必需的,因为如果在网络 IO 线程或业务线程中直接往磁盘写数据的话,写操作偶尔可能阻塞长达数秒之久(原因很复杂,可能是磁盘或磁盘控制器复位)。这可能导致请求方超时, 或者耽误发送心跳消息,在分布式系统中更可能造成多米诺骨牌效应,例如误报死锁引发自动 failover 等。因此,在正常的实时业务处理流程中应该彻底避免磁盘 IO,这在使用 one loop per thread 模型的非阻塞服务端程序中尤为重要,因为线程是复用的,阻塞线程意味着影响多个客户连接。

### Part II: muduo 网络库

#### chpt 6 muduo 网络库简介

* 见【code-reading笔记】--muduo--Usage

* muduo 是静态链接的 C++ 程序库

  * 原因是在分布式系统中正确安全地发布动态库的成本很高,见第 11 章
  * 使用 muduo 库的时候,只需要设置好头文件路径(例如 ../build/debug-install/include)和库文件路径(例如 ../build/debug-install/lib)并链接相应的静态库文件(-lmuduo_net -lmuduo_base)即可

* 介绍了一下目录结构

* 使用教程

  * TCP 网络编程本质论:处理三个半事件

    * 连接的建立,包括服务端接受(accept)新连接和客户端成功发起(connect) 连接。TCP 连接一旦建立,客户端和服务端是平等的,可以各自收发数据。
    * 连接的断开,包括主动断开(close、shutdown)和被动断开(read(2)返回0)。
    * 消息到达,文件描述符可读。这是最为重要的一个事件,对它的处理方式决定 了网络编程的风格(阻塞还是非阻塞,如何处理分包,应用层的缓冲如何设计,等等)。
    * 消息发送完毕,这算半个。对于低流量的服务,可以不必关心这个事件;另外,这里的“发送完毕”是指将数据写入操作系统的缓冲区,将由 TCP 协议栈负责数据的发送与重传,不代表对方已经收到数据。

  * 网络编程难点:

    * 如果要主动关闭连接,如何保证对方已经收到全部数据?如果应用层有缓冲(这 在非阻塞网络编程中是必需的,见下文),那么如何保证先发送完缓冲区中的数据, 然后再断开连接?直接调用 close(2) 恐怕是不行的

    * 如果主动发起连接,但是对方主动拒绝,如何定期(带 back-off 地)重试?

    * 非阻塞网络编程该用边沿触发(edge trigger)还是电平触发(level trigger)? 

      * 如果是电平触发,那么什么时候关注 EPOLLOUT 事件?会不会造成 busy-loop?
      * 如果是边沿触发,如何防止漏读造成的饥饿?epoll(4) 一定比 poll(2) 快吗?

    * 非阻塞网络编程中,为什么要使用应用层发送缓冲区?

    * 在非阻塞网络编程中,为什么要使用应用层接收缓冲区?

      * 假如一次读到的数据不够一个完整的数据包,那么这些已经读到的数据是不是应该先暂存在某个地方, 等剩余的数据收到之后再一并处理?见 [lighttpd 关于 \r\n\r\n 分包的 bug](https://redmine.lighttpd.net/issues/2105)。

      * 假如数据是一个字节一个字节地到达,间隔 10ms,每个字节触发一次文件描述符可读

        (readable)事件,程序是否还能正常工作? lighttpd 在这个问题上出过[安全漏洞](https://download.lighttpd.net/lighttpd/security/lighttpd_sa_2010_01.txt)

    * 非阻塞网络编程中,如何设计并使用缓冲区?

      * muduo 用 readv(2) 结合栈上空间巧妙地解决了这个问题

    * 如果使用发送缓冲区,万一接收方处理缓慢,数据会不会一直堆积在发送方,造
      成内存暴涨?如何做应用层的流量控制?

    * 如何设计并实现定时器? 并使之与网络 IO 共用一个线程,以避免锁

* 性能评测
  * 擅长TCP长连接
  * example/ping_pong 击鼓传花
  * v.s. libevent2:
    * 每次从socket最大读取字节数
    * `epoll_ctl(fd, EPOLL_CTL_ADD, ...) ` 更新event_watcher

* 详解 muduo 多线程模型

  * 协议带上id,以支持parallel pipelining:响应中会回显请求中的 id,client可以不假设response的顺序性
  * 常见的并发网络服务程序设计方案
    * ![scalable-server](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Linux多线程服务端编程-muduo/scalable-server.jpeg)
    * benchmark数据:
      * fork()+exit(): 534.7μs。
      * pthread_create()+pthread_join(): 42.5μs,其中创建线程用了 26.1μs。
      * push/pop a blocking queue : 11.5μs。
      * Sudoku resolve: 100us (根据题目难度不同,浮动范围 20~200μs)。
    * 突发请求 和 顺序性,似乎是一对矛盾
  * 方案0
    * recipes/python/echo-iterative.py
    * 适合 daytime 这种 write-only 短连接服务
  * 方案1:echo-fork.py
    * process-per-connection
    * `ForkingTCPServer`
    * 适合长连接 + “计算响应的工作量远大于 fork() 的开销”
  * 方案2:`ThreadingTCPServer`
  * TCP是全双工通信协议,如何实现?
    * 思路1:两个线程,例子 [python pinhole](https://code.activestate.com/recipes/114642/)
    * 思路2:IO multiplexing,也就是 select/poll/epoll/kqueue 这一 列的“多路选择器”,让一个 thread of control 能处理多个连接。
      * 非阻塞编程、设计应用层buffer(7.4)
      * echo_poll.py
      * 对于 listening fd,接受(accept)新连接,并注册到 IO 事件关注列表(watch list),然后把连接添加到 connections 字典中(L18~L23)
    * Doug Schmidt 指出,其实网络编程中有很多是事务性(routine)的工作,可以提取为公用的框架或库,而用户只需要填上关键的业务逻辑代码,并将回调注册到框架中,就可以实现完整的网络服务,这正是 Reactor 模式的主要思想。
  * 方案5:单线程reactor
    * 比方案2多一次系统调用
    * 不适合CPU密集
    * `server_basic.cc`, `recipe/echo-reactor.py`
    * 注意在使用非阻塞 IO +事件驱动方式编程的时候,一定要注意避免在事件回调中执行耗时的操作,包括阻塞 IO 等,否则会影响程序的响应
  * 方案6:过渡方案
  * 方案7:每个连接固定一个计算线程
    * 相比方案6,有顺序性(连接级别的);对突发请求处理不好
  * 方案8:server_threadpool.cc
    * 另外也可以用线程池来调用一些阻塞的 IO 函数,例如 fsync(2)/fdatasync(2),这两个函数没有非阻塞的版本
      * [fsync() on a different thread: apparently a useless trick](http://oldblog.antirez.com/post/fsync-different-thread-useless.html)
      * fsync: slow(几十ms),redis APPEND ONLY mode(fsync never、fsync everysec、fsync always)
  
  * 方案9:one loop per thread
    * 与方案 8 的线程池相比,方案 9 减少了进出 thread pool 的两次上下文切换,在把多个连接分散到多个 Reactor 线程之后,小规模计算可以在当前 IO 线程完成并发回结果,从而降低响应的延迟
    
    * 优化突发请求考虑方案11
  
    * 框架:muduo、netty
    
  * 方案10:reactors in processes
    * 框架:nginx
    
    * 如果连接之间无交互,这种方案也是很好的选择。
  
    * 工作进程之间相互独立,可以热升级。
    
  * 方案11:混合方案8和方案9
    
  * 关于event loop个数
    * [**What is the optimal number of I/O threads for best performance?**](http://wiki.zeromq.org/area:faq#toc3)
    
    * The basic heuristic is to allocate 1 I/O thread in the context for every gigabit per second of data that will be sent and received (aggregated). Further, the number of I/O threads should **not exceed** (number_of_cpu_cores - 1).
  
    * 如果 TCP 连接有优先级之分,那么单个 event loop 可能不适合,正确的做法是把高优先级的连接用单独的 event loop 来处理
      * 在 muduo 中,属于同一个 event loop 的连接之间没有事件优先级的差别。这么设计的原因是为了防止优先级反转
    
  * 一些web server编程材料
    * https://gee.cs.oswego.edu/dl/cpjslides/nio.pdf
    * http://www.kegel.com/c10k.html
    * http://bulk.fefe.de/scalable-networking.pdf



#### chpt 7 muduo 编程示例

* 五个简单TCP示例
  * echo(RFC 862)、discard(RFC 863)、chargen(RFC 864)、daytime(RFC 867)、time(RFC 868)
  * 以上几个协议的消息格式都非常简单,没有涉及 TCP 网络编程中常见的分包处理,在后文 §7.3 讲 Boost.Asio 的聊天服务器时我们再来讨论这个问题。
* 文件传输
  * 在网络编程中,应用程序发送数据往往比接收数据简单(实现非阻塞网络库正相反,发送比接收难)
* Boost.Asio 的聊天服务器
  * 长连接的分包:长度字段
* muduo Buffer 类的设计与使用
  * muduo 的 IO 模型
    * 在 event handler 中,程序要尽快交出控制权,返回窗口的事件循环
  * 为什么 non-blocking 网络编程中应用层 buffer 是必需的
    * output buffer: 核心是不阻塞带来的约束
    * input buffer: 数据的不完整性
  * muduo EventLoop 采用的是 epoll(4) level trigger,而不是 edge trigger。一是 为了与传统的 poll(2) 兼容,因为在文件描述符数目较少,活动文件描述符比例较高时,epoll(4) 不见得比 poll(2) 更高效,必要时可以在进程启动时切换 Poller。 二是 level trigger 编程更容易,以往 select(2)/poll(2) 的经验都可以继续用,不可能发生漏掉事件的 bug。三是读写的时候不必等候出现 EAGAIN,可以节省系统调用次数,降低延迟
    * [What is the purpose of epoll's edge triggered option?](https://stackoverflow.com/questions/9162712/what-is-the-purpose-of-epolls-edge-triggered-option)  
    * edge trigger 灵活度更高,可以在收到信号后自己决策read/write时机
    * level trigger,用户只需要感知 input/output buffer,不用自己操作read/write
  * Buffer 的功能需求
  * 关于性能
    * 如果确实在内存带宽方面遇到问题,说明你做的应用实在太 critical,或许应该 考虑放到 Linux kernel 里边去,而不是在用户态尝试各种优化。毕竟只有把程序做到 kernel 里才能真正实现 zero copy;否则,核心态和用户态之间始终是有一次内存拷贝的。如果放到 kernel 里还不能满足需求,那么要么自己写新的 kernel,或者直接用 FPGA 或 ASIC 操作 network adapter 来实现你的“高性能服务器”。

* 一种自动反射消息类型的 Protobuf 网络传输方案

  * 网络编程中使用 Protobuf 的两个先决条件

    * 自己处理长度信息+类型信息
    * “山寨”做法:类型信息用唯一的typeid表示或者以name作key用全局的lookup table

  * reflection:根据 type name 反射自动创建 Message 对象

  * Protobuf 传输格式

    * ```c++
      struct ProtobufTransportFormat __attribute__ ((__packed__))
      {
        int32_t  len;
        int32_t  nameLen;
        char     typeName[nameLen];
        char     protobufData[len-nameLen-8];
        int32_t  checkSum; // adler32 of nameLen, typeName and protobufData
      };
* signed int。消息中的长度字段只使用了 signed 32-bit int,而没有使用 unsigned int,这是为了跨语言移植性,因为 Java 语言没有 unsigned 类型。另外,Protobuf 一般用于打包小于 1MB 的数据,unsigned int 也没用。
* check sum。虽然 TCP 是可靠传输协议,虽然 Ethernet 有 CRC-32 校验,但是 网络传输必须要考虑数据损坏的情况,对于关键的网络应用,check sum 是必不可少的。见 § A.1.13 “TCP 的可靠性有多高”。对于 Protobuf 这种紧凑的二进 制格式而言,肉眼看不出数据有没有问题,需要用 check sum。
* adler32 算法。我没有选用常见的 CRC-32,而是选用了 adler32,因为它的计算 量小、速度比较快,强度和 CRC-32 差不多。另外,zlib 和 java.unit.zip 都直接支持这个算法,不用我们自己实现。
* type name 以 '\0' 结束。这是为了方便 troubleshooting,比如通过 tcpdump 抓下来的包可以用肉眼很容易看出 type name,而不用根据 nameLen 去一个个数字节。同时,为了方便接收方处理,加入了 nameLen,节省了 strlen(),这 是以空间换时间的做法。
* 没有版本号。Protobuf Message 的一个突出优点是用 optional fields 来避免协议的版本号
  • 在 muduo 中实现 Protobuf 编解码器与消息分发器
  • 为什么 Protobuf 的默认序列化格式没有包含消息的长度与类型
* rpc/tcp port均能判断消息类型
* `service SudokuService { rpc Solve (SudokuRequest) returns (SudokuResponse);`
* 只有在使用 TCP 长连接,且在一个连接上传递不止一种消息的情况下(比方同时发 Heartbeat 和 Request/Response,见9.3),才需要我前文提到的那种打包方案。这时候我们需要一个分发器 dispatcher,把不同类型的消息分给各个消息处理函数
  • 什么是编解码器(codec): encode+decode
* 传输格式 <-> Buffer
  • example/protobuf/codec*
  • 消息分发器(dispatcher)有什么用
* ProtobufCodec 与 ProtobufDispatcher 的综合运用
* 在构造函数中,通过注册回调函数把四方(TcpConnection、codec、dispatcher、
QueryServer)结合起来
  • ProtobufDispatcher 的两种实现
  • ProtobufCodec 和 ProtobufDispatcher 有何意义
* §9.7 “分布式程序的自动化回归测试”会介绍利用 Protobuf 的跨语言特性,采用 Java 为 C++ 服务程序编写 test harness。
* 这种编码方案的 Java Netty 示例代码见 http://github.com/chenshuo/muduo-protorpc 中的 com.chenshuo.muduo.codec package。
  • 限制服务器的最大并发连接数

Part IV: 附录

网络编程学习经验

  • 计算机网络是个 big topic,涉及很多人物和角色,既有开发人员,也有运维人员。比方说:公司内部两台机器之间 ping 不通,通常由网络运维人员解决,看看是 布线有问题还是路由器设置不对;两台机器能 ping 通,但是程序连不上,经检查是 本机防火墙设置有问题,通常由系统管理员解决;两台机器能连上,但是丢包很严重,发现是网卡或者交换机的网口故障,由硬件维修人员解决;两台机器的程序能连上,但是偶尔发过去的请求得不到响应,通常是程序 bug,应该由开发人员解决。
  • 面向业务的网络编程的特点
    • 不一定需要遵循公认的通信协议标准:如果用短连接 TCP 协议,为了优化性能通常要精心设计 accept 新连接的机制,避免惊群并减少上下文切换。但是如果改用长连接, 用最简单的单线程 accept 就行了
    • 现在的机器上,简单的并发长连接 echo 服务程序不用特别优化就做到十多万 qps,但是如果每个业务请求需要 1ms 密集计算,在 8 核机器上充其量能达到 8 000 qps,优化 IO 不如去优化业务计算(如果投入产出合算的话)。
  • 几个术语
    • 在 TCP 网络编程中,客户端和服务端很容易区分,主动发起连接的是客户端,被动接受连接的是服务端。当然,这个“客户端”本身也可能是个后台服务程序,HTTP proxy 对 HTTP server 来说就是个客户端。
  • 7 × 24 重要吗,内存碎片可怕吗
    • allocator很成熟了
    • 普通 PC 服务器的年故障率约为 3% ~ 5%
  • 协议设计是网络编程的核心
    • 关闭连接。在传统的网络服务中(特别是短连接服务),不少是服务端主动关闭连接,比如 daytime、HTTP 1.0。也有少部分是客户端主动关闭连接,通常是些长连接服务,比如 echo、chargen 等。我们自己的业务系统该如何设计连接关闭协议呢?
* 服务端主动关闭连接的缺点之一是会多占用服务器资源。服务端主动关闭连接之后会进入TCP的TIME_WAIT 状态,在一段时间之内持有(hold)一些内核资源。如果并发访问量很高,就会影响服务端的处理能力。这似乎暗示我们应该把协议设计为客户端主动关闭,让 TIME_WAIT 状态分散到多台客户机器上,化整为零。
  • 消息设计,一个消息应该包含哪些内容?
* 多个程序相互通信如何避免 race condition?(p. 348)
* 外部事件发生时,网络消息应该发 snapshot 还是 delta?
* 新增功能时,各个组件如何平滑升级?
  • end-to-end principle 和 happens-before relationship
  • 网络编程的三个层次
    • 熟悉Linux 的 TCP/IP 协议栈的脾气
* 有可能出现 TCP 自连接(self-connection),程序应该有所准备
  * 见8.11和[《学之者生,用之者死——ACE 历史与简评》](https://blog.csdn.net/solstice/article/details/5364096) 
  * 三个硬伤:sleep < 2ms、Linux TCP self-connection、timeval on 64-bit
* 内核可能有bug
* 写可靠的网络程序的关键是熟悉各种场景下的 error code (文件描述符用完了如何?本地 ephemeral port 暂时用完,不能发起新连接怎么办?服务端新建并发连接太快,backlog 用完了,客户端 connect 会返回什么错误?)
  • 最主要的三个例子
    • echo、chat、proxy
    • proxy 的作用:连接的管理更加复杂:既要被动接受连接,也要主动发起连接; 既要主动关闭连接,也要被动关闭连接。还要考虑两边速度不匹配(§ 7.13)
  • 学习 Sockets API 的利器:IPython
    • 在编写 muduo 的时候,我一般会开四个命令行窗口,其一看 log,其二看 strace, 其三用 netcat/tempest/ipython 充作通信对方,其四看 tcpdump
$ ipython
In [1]: import socket, select
In [2]: s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
In [3]: s.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
In [4]: s.bind(('', 5000))
In [5]: s.listen(5) # backlog queue len == 5
In [6]: client, address = s.accept()  # client.fileno()
In [7]: client.recv(1024) # 此处会阻塞 Out[7]: 'Hello\n'
In [8]: epoll = select.epoll()
In [9]: epoll.register(client.fileno(), select.EPOLLIN) # 试试省略第二个参数
In [10]: epoll.poll(60) # 此处会阻塞
Out[10]: [(4, 1)] # 表示第 4 号文件可读(select.EPOLLIN == 1)
In [11]: client.recv(1024) # 已经有数据可读,不会阻塞了 Out[11]: 'World\n'
In [12]: client.setblocking(0) # 改为非阻塞方式
In [13]: client.recv(1024) # 没有数据可读,立刻返回,错误码 EAGAIN == 11 error: [Errno 11] Resource temporarily unavailable
In [14]: epoll.poll(60) # epoll_wait() 一下 Out[14]: [(4, 1)]
In [15]: client.recv(1024) # 再去读数据,立刻返回结果 Out[15]: 'Bye!\n'
In [16]: client.close()

$ nc localhost 5000
Hello <enter>
World <enter>
Bye! <enter>
  • TCP 的可靠性有多高
  • Realize That TCP Is a Reliable Protocol, Not an Infallible Protocol.
  • IP header 和 TCP header 的 checksum 是一种非常弱的 16-bit check sum 算法
* 比如没法检查出两个16-bit整数的交换
* 以太网的 CRC32 只能保证同一个网段上的通信不会出错(两台机器的网线插到同一个交换机上,这时候以太网的 CRC 是有用的)。但是,如果两台机器之间经过了多级路由器呢?
  * NAT会替换源地址,这是TCP payload无法通过TCP header checksum校验
  • 路由器可能出现硬件故障,比方说它的内存故障(或偶然错误)导致收发 IP 报文出现多 bit 的反转或双字节交换,这个反转如果发生在 payload 区,那么无法用链路层、网络层、传输层的 check sum 查出来,只能通过应用层的 check sum 来检测。
  • 另外一个例证:下载大文件的时候一般都会附上 MD5,这除了有安全方面的考虑(防止篡改),也说明应用层应该自己设法校验数据的正确性。这是 end-to-end principle 的一个例证
  • 相关资料:
* 《When the CRC and TCP checksum disagree》
* [《The Limitations of the Ethernet CRC and TCP/IP checksums for error detection》](http://noahdavids.org/self_published/CRC_and_checksum.html)
  * 1 in 16 million and 1 in 10 billion TCP segments
  * 33.91 hours on a gigabit network
  * 建议:zip、md5
  • 书籍推荐
  • 《TCP/IP Illustrated, Vol. 1: The Protocols》TCPv1
  • 《Unix Network Programming, Vol. 1: Networking API》UNP
  • 《Effective TCP/IP Programming》
  • 《TCP/IP Illustrated, Vol. 2: The Implementation》TCPv2
  • 《Pattern-Oriented Software Architecture Volume 2: Patterns for Concurrent
and Networked Objects》POSA2

操作系统

操作系统——本科课程

[toc]

书1:现代操作系统,答案

chpt1 引论

  • 工作原理和实现方法
  • 技术含量最高的系统软件

001

  • 将硬件的复杂性与程序员分离开
  • 定义:系统软件, 程序模块的集合,资源管理和用户接口功能

作用:

  1. 作为扩展机器
  2. 裸机添加
  3. 理解“抽象”,如文件
  4. 作为资源管理者/作业管理(大型机)
  5. 多路复用,时间与空间
  6. 软硬件接口
  7. system call

1.发展史

  1. 批处理操作系统
改进主存和I/O设备之间的吞吐量

特征:用户脱机使用,无法交互
  1. 多道程序设计 multiprogramming
需要硬件保护

宏观上并行,微观上串行
  1. 系列机思想与IBM System/360系统
统一的体系结构、指令集 -> 软件危机
  1. 分时系统:轮流使用time slice
unix操作系统:多用户、多任务、分时

unix前身是Multics系统,内外环设计
  1. 大规模集成电路时代
分布式、嵌入式

2.操作系统的硬件环境

寄存器分类:

  • 用户可见寄存器: 高级语言编译器通过算法分配并使用之,以减少程序访问主存次数
    • 数据寄存器、地址寄存器、条件码寄存器(溢出、符号)
  • 控制和状态寄存器: 用于控制处理器的操作,由OS的特权代码使用, 以控制其他程序的执行
    • PC、IR、PSW(program status word)
    • PSW中有一位CPU的状态工作码, PSW在系统调用和I/O中很重要

核心态=管态、用户态=目态

  • I/O不在用户态内
  • 用户态->核心态: trap 核心态->用户态:修改PSW
  • e.g. x86: R0 R1 R2 R3

层次化存储体系结构 ~ 容量、速度、成本

存储访问局部性原理
cpu访问寄存器不存在延时,操作系统不访问cache
磁带光盘价格极其低廉,作为磁盘的备份

shell 、GUI
多线程和多核芯片 书1p13

I/O:控制器+设备本身 device driver
002

  • 实现I/O的三种方式

轮询、中断、DMA芯片(总线竞争、大量I/O数据传送)

中断系统的两大组成部分: 硬件中断装置和软件中断处理程序 (中断设备的设备驱动程序的一部分)

中断向量 中断向量表IVT

005

启动计算机:主板BIOS
进程、地址空间、进程间通信
虚拟内存
UNIX:特殊文件、管道
rwx r-x --x p25


**“个体重复系统发育”**,很深刻

3.系统调用

1.6 系统调用~执行“trap” 功能号和参数
大多采用在内存中开辟专用堆栈区来传递参数

linux系统调用:
system_call() sys_call_table

Linux系统调用利用了x86体系结构中的软件中断,即调用了int $0x80汇编指令。这条汇编指令将产生向量为128的软件中断, CPU被切换到内核态,并将控制权交给系统调用过程的起点system_call()处理函数

系统调用与内核函数(即服务例程,e.g. sys_getpid())
封装例程 ppt2-p56

006

用于进程管理、文件管理、目录管理 书1p32
UID、GID、PID
PID~fork,waitpid指令
fork不可继承的有:进程标识符,父进程标识符
execve:替换一个进程的核心映像

POSIX过程调用 :进程管理(fork, execve, waitpid, exit(status) )、文件管理(open, ...)、目录管理(mkdir, rmdir, ..., umount, unlink)、其它(chmod, kill, chdir, time)

<->

WIN32 API:

003

4.操作系统结构

  • 单体系统(模块组合结构)
  • 层次式系统
    • 分层:全序、偏序、半序
  • 虚拟机结构: VM/370

会话监控系统CMS
004 这是1型超级监控程序(cons:用户态不能陷入)

2型:VMWare,主机/客户操作系统

  • 微内核结构:
    • 运行在核心态的内核􏰁提供最基本的操作系统功能,包括中断处理、处理机调度、进程间通信。这些部分只􏰁供了一个很小的功能集合,通常称为微内核
    • 特点:Mechanism和policy分离=>使内核更小
    • 微内核结构的变体:客户-服务器模型
    • 客户进程与服务器进程之间使用消息进行通信

Windows内核结构

chpt2 进程与线程

1.进程

顺序进程模型

  • 有时须考虑严格的实时要求

并发:次序不是事先确定的

进程是具有独立功能的程序在某个数据集合上的一次运行活动,是系统进行资源分配和调度的独立单位

  • 资源分组处理与执行
  • 进程的组成:程序+数据+PCB进程控制块+堆栈
  • 进程上下文
    • 用户级上下文:进程的用户地址空间(包括用户栈各层次),包括用户正文(code)段、用户数据段和用户栈;
    • 寄存器级上下文:程序寄存器、处理机状态寄存器、栈指针、通用寄存器的值;
    • 系统级上下文:
* 静态部分(PCB和资源表格)
* 动态部分:核心栈 (核心过程的栈结构,不同进程在调用相同核心过程时有不同核心栈)
  • 注意有两种register saves/restores:
    • timer interrupt: 用hardware,kernel stack,implicitly,存user registers
    • OS switch:用software,process structure,explicitly,存kernel registers

NOTE:

  • 守护进程daemon
  • windows没有进程层次的概念 书p51
  • shell p26

调度进程: 进程表=PCB表 p53

  • 并发度:PCB表的大小
  • 链表结构、索引结构
  • 多道(内存层面)!=系统并发度(OS层面)

012

  • windows进程无状态,只是宿主,真正运行的是线程(调度单位)
  • UNIX中,OS作为进程的一部分
e.g. Linux进程控制块
  • struct task_struct 双向循环链表
  • 调度针对就绪进程 run_list
进程的状态
008
  • 运行->阻塞:OS、I/O、资源、IPC(其它进程输入)
  • 运行->就绪:时间片用完、因为高优先级进程就绪而中断
  • 就绪->运行:调度程序
  • 阻塞->就绪

其它:

  • initial和final(五状态):
    • initial态:等待资源;
    • final态(在UNIX称作zombie state)等待子进程return 0,parent进程 wait()子进程
    • 表格和其它信息暂时由辅助程序保留,例如为处理用户帐单而累计资源使用情况的财务

程序

  • suspend(挂起)状态(七状态):进程映像在磁盘上,不占用内存空间
  • 激活的概念:阻塞挂起->阻塞, 就绪挂起->就绪
阻塞和唤醒 linux进程状态

2.进程控制

  • 进程控制原语:创建、撤销、阻塞、唤醒
    • 系统调用不一定是原语:两进程,read()重入
e.g. POSIX进程控制
  • 创建进程
    • fork——创建新进程,复制现有进程上下文
    • exec——加载新程序并覆盖自身
  • 可继承(inherit):子进程可以从父进程中继承用户标识符、环境变量、打开文件、 文件系统的当前目录、控制终端、已经连接的共享存储区、信号处理例程入口表等
  • 不可继承:进程标识符,父进程标识符
  • exit:exit()向父进程给出一个退出码(8位的整数),父进程终止时如何影响子进程:
    • 子进程从父进程继承了进程组ID和终端组ID(控制终端),因此子进程对发给该进程组或终端组的信号敏感。终端关闭时,以该终端为控制终端的所有进程都收到SIGHUP信号。
    • 子进程终止时,向父进程发送SIGCHLD信号,父进程截获此信号并通过wait()系统调用来释放子进程PCB
  • wait() waitpid() waitid()等待子进程修改状态
e.g. Linux进程控制
#include <cstdlib>
#include <iostream>
#include <sys/types.h>
#include<sys/wait.h>
#include <unistd.h>
#include <vector>
using namespace std;

int main(){
    int pid;
    for (int i = 0; i < 3; i++){
        pid = fork();
        if (pid == 0){ //子进程
            cout << "pid=" << pid << getpid() << ", hello world" << endl;
            return 0;
        }
        else{
            cout << "pid=" << pid << " forked" << endl;
        }
    }
    int ret = 0;
    while (ret!=-1){
        ret = wait(NULL);
        cout << "pid=" << ret << " exited" << endl;
    };
    return 0;
}
e.g. Windows进程控制
/*
BOOL CreateProcess(
LPCWSTR pszImageName, LPCWSTR pszCmdLine, LPSECURITY_ATTRIBUTES psaProcess, LPSECURITY_ATTRIBUTES psaThread, BOOL fInheritHandles,
DWORD fdwCreate, LPVOID pvEnvironment, LPWSTR pszCurDir,
LPSTARTUPINFOW psiStartInfo, LPPROCESS_INFORMATION pProcInfo
);
*/
#include <stdio.h> 
#include <windows.h>
int main() {
	TCHAR szCmdLine[]={TEXT("C:\\oscourse\\test.exe")}; 
	STARTUPINFO si;
	PROCESS_INFORMATION pi;
	memset(&si, 0, sizeof(STARTUPINFO));
	si.cb = sizeof(STARTUPINFO); 
  si.dwFlags = STARTF_USESHOWWINDOW; 
  si.wShowWindow = SW_SHOW;
  if(!CreateProcess(NULL, szCmdLine, NULL, NULL, FALSE, 0, NULL, NULL, &si, &pi))
	{
		printf("Create process fail!\n");
		ExitProcess(1);
	} 
  else {
    printf("Create process success!\n"); 
    ExitProcess(0);
	} 
}

3.线程模型

进程是资源分配单位 ~ 并发执行

  • 1-p^n 内存与吞吐量

线程是CPU调度单位 ~ 时空开销

  • 多线程Multithreading
015 016
  • 线程:并行实体共享同一地址空间和所有可用数据,线程比进程更容易创建和撤销
  • 轻量级进程=Lightweight process=LWP
  • TCB与PCB
  • e.g. 书p56 三线程、文字处理
  • 在支持线程的操作系统中,进程只作为资源分配单位,而线程则作为CPU调度单位,可读取全局变量来实现通信,比进程间通信简单不少,无需调用内核

NOTE:

  • 服务器与cache,分派线程、工作线程,c代码
  • 线程改善了web服务器的性能
  • 有限状态机实现,非阻塞系统调用<->顺序进程模型的保留 书p57
  • [“阻塞非阻塞"与"同步异步”]( https://www.cnblogs.com/skying555/p/5028167.html): 线程可以异步,因为有状态共享;同步是两个对象之间的关系,而阻塞是一个对象的状态。

4.线程的实现机制

在用户空间/内核中实现线程

用户级线程(ULT)
  • 线程库(运行时系统 run-time system):POSIX的pthread
  • pros:
    • 线程切换不调用操作系统内核,性能良好
    • 调度是应用程序特定的,可针对应用优化
    • ULT可运行在任何操作系统上 (仅需线程库)
  • cons:
    • 调度通常采用非抢先式和更简单的规则(线程没有时间中断=>只能用非抢占式)
    • 阻塞系统、缺页中断,进程中所有线程将被阻塞
    • 操作系统内核只将处理器分配给进程,同一进程中的两个线程不能同时运行于两个处理器上
内核级线程(KLT)
  • pros:
    • 对于多处理器,内核可以同时调度同一进程的多个线程
    • 阻塞是在线程一级完成
    • 内核例程是多线程的
  • cons:
    • 线程管理代价大:在同一进程内的线程切换调用内核,导致速度下降
    • 时间片分配给线程,所以多线程的进程获得更多CPU时间
混合实现

Solaris

在线程的混合实现机制中,线程创建、调度、 同步在用户级线程中完成,应用程序的多个ULT将被映射到一些KLT上

4.Windows的线程

  • KLT
  • 惰性进程,必须有一个线程(自动创建主线程)
  • Windows线程由执行体线程块ETHREAD表示,即线程对象,其中包含内核线程块 KTHREAD,即线程控制块TCB
#include <windows.h>
#include <stdio.h>
#define MAX_THREADS 3
typedef struct _MyData {
	int val1;
	int val2;
} MYDATA, *PMYDATA;

DWORD WINAPI ThreadProc(LPVOID lpParam){
	PMYDATA pData;
	pData = (PMYDATA)lpParam;
	printf("This is thread %d, the parameter is %d\n", pData->val1, pData->val2);
	// Free the memory allocated by the caller for the thread
	// data structure.
	HeapFree(GetProcessHeap(), 0, pData);
	return 0; 
}

void main() {
	PMYDATA pData;
	DWORD dwThreadId[MAX_THREADS]; //线程ID
	HANDLE hThread[MAX_THREADS]; //线程句柄 int i;
	// Create MAX_THREADS worker threads.
	for( i=0; i<MAX_THREADS; i++ ) {
	// Allocate memory for thread data.
		pData = HeapAlloc(GetProcessHeap(), HEAP_ZERO_MEMORY, sizeof(MYDATA));
    if( pData == NULL ) ExitProcess(2);
    // Generate unique data for each thread.
    pData->val1 = i; 
    pData->val2 = i+100; 
    hThread[i] = CreateThread(
                NULL,              // default security attributes
                0,                 // use default stack size
                ThreadProc,        // thread function
                pData,             // argument to thread function
                0,                 // use default creation flags
                &dwThreadId[i]);   // returns the thread identifier

    // Check the return value for success.
    if (hThread[i] == NULL) {
      ExitProcess(i);
    }
	}
// Wait until all threads have terminated. WaitForMultipleObjects(MAX_THREADS, hThread, TRUE, INFINITE); // Close all thread handles upon completion.
	for(i=0; i<MAX_THREADS; i++) {
		CloseHandle(hThread[i]);
	}
}

5.POSIX线程

内核和用户空间实现均可

pthread调用
  • pthread_yield阻塞线程
#define _REENTRANT
#include <pthread.h>
#include <stdio.h>
#include <unistd.h>
#define NUM_THREADS 5
#define SLEEP_TIME 10
void *sleeping(void *); /* thread routine */
int i;
pthread_t tid[NUM_THREADS]; /* array of thread IDs */

int main()
{
    int sleep_time = SLEEP_TIME;
    for (i = 0; i < NUM_THREADS; i++)
        pthread_create(&tid[i], NULL, sleeping, (void *)&sleep_time);
    for (i = 0; i < NUM_THREADS; i++)
        pthread_join(tid[i], NULL); //逐个等待线程结束
    printf("main() reporting that all %d threads have terminated\n", i);
    return (0);
} /* main */

void *sleeping(void* arg)
{
    int sleep_time = *(int *)arg;
    printf("thread %u sleeping %d seconds ...\n", pthread_self(), sleep_time);
    sleep(sleep_time);
    printf("\nthread %u awakening\n", pthread_self());
    pthread_exit(0);
}

6.Linux的线程

  • 从内核的角度,Linux只有进程而没有线程的概念,Linux没有准备特别的调度算法或定义特别的数据结构来表征线程
  • 在Linux中,线程仅仅被视为使用某些共享资源的进程,每个线程都拥有自己的task_struct,所以在内核看来它就是一个普通的进程,但是该进程与其他进程共享某些资源,例如地址空间
  • Linux下的pthread_create是通过clone系统调用实现的
017
#define _GNU_SOURCE
#include <stdio.h> 
#include <errno.h> 
#include <sched.h> 
#include <sys/types.h>
#define STACK_SIZE 4096 //4k 
int flag;
void *test(void *arg)
{
    int childnum;
  	flag = 1;
    childnum = *(int *)arg;
    printf("Thread %d work cycle\n", childnum);
    sleep(3);
}
int main()
{
    pid_t pid;
    int childno = 1, mainnum = 0;
    void *csp, *tcsp;
    csp = (char *)malloc(STACK_SIZE);
    if (csp)
    {
        tcsp = csp + STACK_SIZE;
    }
    else
    {
        exit(errno);
    }
    flag = 0;
    childno = 1;
    if ((pid = clone((void *)&test, tcsp, CLONE_VM, (void *)&childno)) < 0)
    {
        printf("Couldn't create new thread!\n");
        exit(1);
    }
    else
    { //we're in main
        while (flag == 0); //利用全局变量进行Linux线程间通信
        printf("Just created thread %d\n", pid);
    }
    test(&mainnum);
    sleep(3);
    printf("Main program is now shutting down\n\n");
    return 0;
}

7.进程间通信(IPC)问题概述

进程间关系:互斥、同步、通信

  • 同步是通过共享协作

概念:

  • 竞争条件 (race condition)
  • 临界资源 (critical resource)
  • 临界区(critical section): 进程中访问临界资源的代码片断
  • 进入区->临界区->退出区->剩余区(remainder section)

目的: 避免竞争条件,同时保证使用临界资源的进程能够正确和高效地进行协作

解决互斥问题应遵循的条件

  1. 任何两个进程不能同时处于临界区
  2. 不应对CPU的速度和数量做任何假设
  3. 临界区外运行的进程不得阻塞其他进程
  4. 不得使进程无限期等待进入临界区

8.互斥问题的几种算法

禁止中断
  • 简单
  • 把禁止中断的权利交给用户进程导致系统可靠性较差
  • 不适用于多处理器(违反条件2)
  • 可能错过其它interrupts,比如disk的read request
  • 对于现代CPU,这个实现速度慢
共享锁变量
while(lock==1);
lock=1;
//critical region
lock=0;
//non_critical region;
  • 可能违反条件1
  • 忙等待(自旋锁spin lock)
  • Murphy's law
严格轮转法
  • 整型变量turn
  • 可能违反条件3(该算法要求两个进程严格轮流进入临界区)
  • 忙等待
  • e.g. spooler directory 假脱机程序
Peterson算法

丢失则进入临界区; 解决了互斥访问的问题,而且克服了强制轮转法的缺点,可以正常地工作.

忙等待

#define FALSE 0
#define TRUE  1
#define N     2                                           /* 进程数量 */
  
int turn=0;                                 /* 现在轮到谁?*/
int interested[N];               /* 所有值初始化为0(FALSE)*/

void enter_region(int process)               /* 进程是0或1 */
{
	interested[process] = TRUE;              /* 表名所感兴趣的*/
	turn = 1-process;                            /* 设置标志 */
	while(turn == 1-process && interested[1-process] ==TRUE); /* 空语句 */
}

void leave_region(int process)                    
{
	interested[process] = FALSE;             /* 表示离开临界区*/ 
}
硬件支持方法
  • test and set lock:X86的BTS/BTR命令
  • swap:XCHG OP1 OP2
  • 优点:
    • 适用于任意数目的进程
    • 简单,容易验证其正确性
    • 可以支持进程中存在多个临界区,只需为每个临界区设立一个布尔变量
  • 缺点:忙等待,耗费CPU时间
优先级反转问题

priority inversion problem:两个进程H、L,H优先级高于L,调度规则: 只要H处于就绪态它就可以运行。若某一时刻L处于临界区中,此时H变成就绪态并被调度,从而开始忙等待;但是由于H的优先级高于L,使得L不会被调度,也就无法离开临界区

9.信号量

  • 信号量是OS提供的管理共享资源的有效手段
  • s.count(), s.queue(),OS核心代码
  • dijkstra:信号量、原子操作
  • 原语P、V来自荷兰语的proberen(降低)和verhogen(升起)
  • 利用信号量实现互斥(mutex)、进程间同步($$S_{12}$$)
    • 同步P操作在互斥P操作前
  • 信号量的缺点:同步操作分散、易读性差、不利于修改和维护、正确性难以保证
  • mutex互斥锁适用于用户级线程包; 快速用户区互斥量futex,书p76
  • 生产者消费者问题
empty=n; 				full=0;
\\producer			\\consumer
P(empty);			  P(full);
P(mutex);			  P(mutex);
...
V(mutex);		    V(mutex);
V(full);				V(empty);

分析:有两个层面的进程间通信,一是生产者和生产者的互斥,二是生产者和消费者的同步,所以必须要有两个信号量。

10.管程

引入
  • 1973年,由Hoare和Hansen提出
  • 管程的基本思想是把信号量及其操作原语封装在一个对象内部。即:将共享变量以及对共享变量能够进行的所有操作集中在一个模块中,解决临界区分散所带来的管理和控制问题
  • 管程的定义:管程是关于共享资源的数据结构及一组针对该资源的操作过程所构成的软件模块
  • 引入管程可提高代码的可读性,便于修改和维护,正确性易于保证
主要特性
  • 模块化:一个管程是一个基本程序单位,可以单独编译
  • 抽象数据类型:管程是一种特殊的数据类型, 其中不仅有数据,而且有对数据进行操作的代码
  • 信息封装:管程是半透明的,进程可调用管程中实现了某些功能(函数) ,至于这些功能是怎样实现的,在其外部则是不可见的
其它
  • 共享变量外部不可见
  • 互斥进入(为了数据完整性)
  • 等待问题:可以规定唤醒是最后操作
  • queue,等待和唤醒操作;入口等待序列,紧急等待队列(多个等待进程)
  • 条件变量:与互斥锁不同,条件变量是用来等待而不是用来上锁的
缺点
  • 管程是一个编程语言概念,编译器必须要识别管程并用某种方式对其互斥做出安排
  • Java支持管程。但是,C及多数程序设计语言都不支持管程
  • C及多数语言也没有信号量,但是增加信号量十分容易——编译器甚至不需知道它们的存在,系统只需提供P、V系统调用即可

11.消息传递

  • 针对分布式系统的高级机制
  • 消息传递原语:send(destination, &message)和receive(source, &message),它们是系统调用而不是语言成分
#define N 100								/* buffer 中槽的数量 */

void producer(void){
  int item;
  message m;									/* buffer 中槽的数量 */
  while(TRUE){
    item = produce_item();						/* 生成放入缓冲区的数据 */
    receive(consumer,&m);						/* 等待消费者发送空缓冲区 */
    build_message(&m,item);						/* 建立一个待发送的消息 */
    send(consumer,&m);								/* 发送给消费者 */
  }
}

void consumer(void){
  int item,i;
  message m;
  for(int i = 0;i < N;i++){						        /* 循环N次 */
    send(producer,&m);							/* 发送N个缓冲区 */
  }
  while(TRUE){
    receive(producer,&m);						/* 接受包含数据的消息 */
  	item = extract_item(&m);					/* 将数据从消息中提取出来 */
    send(producer,&m);							/* 将空缓冲区发送回生产者 */
    consume_item(item);						/* 处理数据 */
  }
}
  • 消息传递系统的设计问题
  • 不可靠消息传递的成功通信问题(网络)
  • 进程命名问题
  • 身份认证问题
  • 性能问题

12.经典IPC问题

哲学家进餐问题
哲学家进餐问题
读者-写者问题
semaphore rmutex = 1; // 读进程互斥信号量
semaphore wmutex = 1; // 写进程互斥信号量
int readcount = 0;    // 读进程计数
semaphore S = 1;      // S的意义在于保证读和写能够公平竞争

cobegin
process reader_i() {  // 读进程
  while(true) {
    P(S);    
    P(rmutex);
    readcount++;
    if(readcount == 1) P(wmutex);
    V(rmutex);
    V(S);
		read_data_base();
    P(rmutex);
    readcount--;
    if(readcount == 0)V(wmutex);
    V(rmutex);
  }
}

process writer_i() {  // 写进程
  while(true) {
    P(S);
    P(wmutex);
    write_data_base();
    V(wmutex);
    V(S);
  }
}
睡眠理发师问题
#define CHAIRS 5 /*为等待的顾客准备的椅子数*/

typedef int semaphone;   /*运用你的想象力*/
semaphore customers = 0; /*等待服务的顾客数*/
semaphore barbers = 0;   /*等待顾客的理发师数*/
semaphore mutex = 1;     /*用于互斥*/
int waiting = 0;         /*等待的顾客(还没理发的)*/

void barber(void)
{
    while (TRUE)
    {
        down(customers);
        /*如果顾客数是0,则睡眠*/
        down(mutex);           /*要求进程等待*/
        waiting = waiting - 1; /*等待顾客数减1*/
        up(barbers);
        /*一个理发师现在开始理发了*/
        up(mutex);  /*释放等待*/
        cut_hair(); /*理发(非临界区操作)*/
    }
}

void customers(void)
{
    down(mutex); /*进入临界区*/
    if (waiting < CHAIRS)
    {                          /*如果没有空椅子,就离开*/
        waiting = waiting + 1; /*等待顾客数加1*/
        up(customers);         /*如果必要的话,唤醒理发师*/
        up(mutex);             /*释放访问等待*/
        down(barbers);         /*如果barbers为0,就入睡*/
        get_haircut();         /*坐下等待服务*/
    }
    else
        up(mutex); /*店里人满了,走吧*/
}

13.Windows的互斥和同步机制

Windows支持三种同步对象

  • 互斥对象(Mutex)
  • 信号量对象(Semaphore)
  • 事件对象(Event)
  • 同步对象等待:
DWORD WaitForSingleObject(
	HANDLE hHandle, // handle of object to wait for 
	DWORD dwMilliseconds // time-out interval in milliseconds
);

DWORD WaitForMultipleObjects(
	DWORD nCount, //对象句柄数组中的句柄数; 
	CONST HANDLE *lpHandles, // 指向对象句柄数组的指针,数组中可包括多种对象句柄;
	BOOL bWaitAll, // 等待标志:TRUE表示所有对象同时可用,FALSE表示至少一个对象可用;
	DWORD dwMilliseconds // 等待超时时限;
);
  • 其它同步方法:
    • 临界区对象(Critical Section):只能在同一进程内使用的临界区,同一进程内各线程对它的访问是互斥进行的
    • 互锁变量访问:相当于硬件指令,对一个整数(进程内的变量或进程间的共享变量)进行操作。其目的是避免线程间切换的影响

14.POSIX的互斥和同步机制

互斥锁
  • 初始化
    • 调用pthread_mutex_init()
    • 或者静态赋值pthread_mutex_t mutex=PTHREAD_MUTEX_INITIALIER
  • 加锁: lock()——阻塞等待锁
  • 解锁: unlock()
  • 清除锁:destroy()——此时锁必需unlock,否则返回EBUSY
条件变量
  • 用来等待而非上锁
  • 函数:pthread_cond_init/wait/timewait/destroy/signal/broadcast
信号量
  • #include
  • 有名(named)信号量和无名(memory-based, unnamed)信号量
  • 无名信号量:
    • 创建:sem_t sem_id;
    • sem_init/getvalue/wait/post/destroy
  • 有名信号量:
    • 特点是把信号量的值保存在文件中。这决定了它的用途非常广:既可以用于线程,也可以用于相关进程间,甚至是不相关进程
    • 有名信号量在使用的时候,和无名信号量共享 sem_wait和sem_post函数;两者的区别是有名信号量使用sem_open代替sem_init,另外在结束的时候要像关闭文件一样去关闭这个有名信号量

14.Linux的进程间通信

Unix早期的进程间通信机制:信号和管道

信号:通知异步事件,软中断
  • e.g. 键盘中断,硬件条件(浮点溢出,segmeatation fault),软件条件(如Socket中有加急数据到达),Shell向子进程发送作业控制命令
  • linux与进程控制相关的命令:&, ctrl-z, fg, bg, ...
  • 进程也可以忽略指定的信号(SIG_IGN), SIGKILL信号(无条件终止进程)和SIGSTOP(使进程暂停)不能被忽略(不能被相关的系统调用阻塞)
  • 由内核执行与该信号相关的默认处理例程(SIG_DFL)
  • 信号的实现:整型,一个字,32位,32种信号,信号是1-index(SIGINT编号是1)
  • task_struct:利用两个字分别记录当前未决的信号(signal)以及当前阻塞的信号(blocked),917行,sigaction存储处理方式
/* Signal handlers: */
	struct signal_struct		*signal;
	struct sighand_struct __rcu		*sighand;
	sigset_t			blocked;
	sigset_t			real_blocked;
	/* Restored if set_restore_sigmask() was used: */
	sigset_t			saved_sigmask;
	struct sigpending		pending;
	unsigned long			sas_ss_sp;
	size_t				sas_ss_size;
	unsigned int			sas_ss_flags;

捕捉signal的实例(shell.md里也有):

#include <signal.h>
void catchint(int signo) {
	printf("\n CATCHINT; signo=%d;", signo); 
	printf("CATCHINT returning\n");
}
void main() {
  int i;
  signal(SIGINT, catchint);
  for(i=0; i<5; i++) {
    // 修改sigaction结构
    printf("Sleep call #%d\n", i);
    sleep(5); 
  }
  printf("Exiting.\n");
}
管道(pipe)
  • 管道类型:pipe、named pipe(还允许无亲缘关系进程通信)
  • 适合数据量大的情况
  • 通过将两个file结构指向同一个临时的VFS索引节点,而 VFS索引节点又指向同一个物理页而实现管道
  • 内核必须利用一定的机制同步对管道的访问, 为此,内核使用了锁、等待队列和信号
  • int pipe(int fildes[2]); fildes[0]是读端,fildes[1]是写端
  • 只适用于父子进程之间;或父进程安排的各个子进程之间 (其它情况用命名管道)
#include <stdio.h> 
#include <unistd.h> 
#include <stdlib.h> 
#include <string.h>
int main(int argc, char *argv[])
{
    int f_des[2];
    static char message[BUFSIZ];
    if (argc != 2)
    {
        fprintf(stderr, "Usage: %s message\n", *argv);
        exit(1);
    }
    if (pipe(f_des) == -1)
    {
        perror("pipe");
        exit(2);
    }
    switch (fork())
    {
        case -1:
            perror("fork");
            exit(3);
        case 0: // 子进程
            close(f_des[1]);
            if (read(f_des[0], message, BUFSIZ) != -1)
            {
                printf("Message received by child:[%s]\n", message);
                fflush(stdout);
            }
            else
            {
                perror("read");
                exit(4);
            }
            break;
        default: // 父进程
            close(f_des[0]);
            strcpy(message, argv[1]);
            if (write(f_des[1], message, BUFSIZ) != -1)
            {
                printf("Message sent by parent:[%s]\n", argv[1]);
                fflush(stdout);
            }
            else
            {
                perror("write");
                exit(5);
            }
    }
}
  • 命名管道:FIFO,和匿名管道的区别在于它是文件实体而不是匿名对象
    • 命名管道可通过mknod系统调用建立:指定mode为S_IFIFO
    • int mknod(const char *path, mode_t mode, dev_t dev);
UNIX System V:消息队列、信号量、共享内存
  • IPC对象(访问必须经过类似文件访问的许可检验)、引用标识符、访问键(定位引用标识符)
    • ipc_perm结构:包含了作为对象所有者和创建者的进程之用户标识符和组标识符,以及对象的 访问模式和对象的访问键。
  • 消息队列
    • 客户/服务器模型,微内核结构,克服了信号承载信息量少,管道只能承载无格式字节流以及缓冲区大小受限等缺点
    • Linux 为系统中所有的消息队列维护一个msgque链表,该链表中的每个指针指向一个msgid_ds结构,该结构完整描述一个消息队列。当建立一个消息队列时,系统从内存中分配一个msgid_ds结构并将指针添加到msgque链表
    • 与消息队列相关的系统调用
* msgget——依据用户给出的键值key,创建新消息队列或打开现有消息队列,返回一个消息队列ID
* msgsnd——发送消息;
	* msgrcv——接收消息,可以指定消息类型;没有消息时,返回-1
	* msgctl——对消息队列进行控制,如删除消息队列
* 消息队列不随创建它的进程的终止而自动撤销,须调用 msgctl(msgqid, IPC_RMID, 0)
  • snd.c
#include <stdlib.h> 
#include <stdio.h> 
#include <string.h> 
#include <errno.h> 
#include <unistd.h> 
#include <sys/msg.h> 
#define MAX_TEXT 512 
#define TRUE 1
struct msgbuf
{
    long int msgtype;
    char msgtext[MAX_TEXT];
};
int main()
{
    struct msgbuf msgdata;
    int msgid;
    char buffer[MAX_TEXT];
    if ((msgid = msgget((key_t)1234, 0666 | IPC_CREAT)) == -1)
    {
        fprintf(stderr, "msgget failed with error: %d\n", errno);
        exit(EXIT_FAILURE);
    }
    printf("msgid = %d\n", msgid);
    while (TRUE)
    {
        printf("Enter message text: ");
        fgets(buffer, MAX_TEXT, stdin);
        msgdata.msgtype = 1;
        strcpy(msgdata.msgtext, buffer);
        if (msgsnd(msgid, (void *)&msgdata, MAX_TEXT, 0) == -1)
        {
            fprintf(stderr, "msgsnd failed\n");
            exit(EXIT_FAILURE);
        }
        if (strncmp(buffer, "end", 3) == 0)
        {
            break;
        }
    }
    exit(EXIT_SUCCESS);
}

rcv.c

#include <stdlib.h> 
#include <stdio.h> 
#include <string.h> 
#include <errno.h> 
#include <unistd.h> 
#include <sys/msg.h> 
#define MAX_TEXT 512 
#define TRUE 1

struct msgbuf
{
    long int msgtype;
    char msgtext[MAX_TEXT];
};
int main()
{
    int msgid;
    struct msgbuf msgdata;
    if ((msgid = msgget((key_t)1234, 0666 | IPC_CREAT)) == -1)
    {
        fprintf(stderr, "msgget failed with error: %d\n", errno);
        exit(EXIT_FAILURE);
    }
    printf("msgid = %d\n", msgid);
    while (TRUE)
    {
        if (msgrcv(msgid, (void *)&msgdata, MAX_TEXT, 0, 0) == -1)
        {
            fprintf(stderr, "msgrcv failed with error: %d\n", errno);
            exit(EXIT_FAILURE);
        }
        printf("Received message: %s", msgdata.msgtext);
        if (strncmp(msgdata.msgtext, "end", 3) == 0)
        {
            break;
        }
    }
    if (msgctl(msgid, IPC_RMID, 0) == -1)
    {
        fprintf(stderr, "msgctl(IPC_RMID) failed\n");
        exit(EXIT_FAILURE);
    }
    exit(EXIT_SUCCESS);
}
  • 信号量
    • semid_ds 结构表示System V IPC信号量
    • sem_base:信号量数组;系统调用参数:信号量索引、操作值和操作标志
    • 操作:semget, semop, semctl
    • 如果系统调用中指定的所有操作中有一个操作不能成功 时,则 Linux会挂起这一进程。但是,如果操作标志指定这种情况下不能挂起进程的话,系统调用返回并指明 信号量上的操作没有成功,而进程可以继续执行
    • 如果进程被挂起,Linux必须保存信号量的操作状态并 将当前进程放入等待队列。为此,Linux在堆栈中建立一个sem_queue结构并填充该结构。新的sem_queue结构添加到信号量对象的等待队列中(利用 sem_pending 和sem_pending_last指针)。当前进程放入sem_queue结构的等待队列中(sleeper)后调用调度程序选择其他的进程运行
semid_ds结构
  • 共享内存
    • 对共享内存的访问同步需要由其他 IPC机制,例如信号量来实现
    • 访问键,访问权限,锁定到物理内存
    • 系统调用
* shmget——创建或打开共享内存:依据用户给出的整数值key,创建新内存区或打开现有内存区,返回 一个共享内存ID
	* shmat——连接共享内存:连接共享内存到本进程的地址空间,可以指定虚拟地址或由系统分配,返回共享内存首地址。父进程已连接的共享内存可被fork 创建的子进程继承
	* shmdt——拆除共享内存连接:拆除共享内存与本进 程地址空间的连接
	* shmctl——共享内存控制:对共享内存进行控制。如共享内存的删除需要显式调用shmctl(shmid, IPC_RMID, 0)
  • 套接字(Socket)
    • 套接字(Socket)是一种网络通信机制,它通过网 络在不同计算机上的进程间进行双向通信。套接字所采用的数据格式可为可靠的字节流或不可靠的报文,通信模式可为client-server模式或 peer-to-peer模式
    • UNIX套接字API(基于TCP/IP):send, sendto, recv, recvfrom

15.Windows的进程间通信

  • 信号量、互斥量、临界区
  • 共享内存:文件映射机制
    • CreateFileMapping/OpenFileMapping
    • MapViewOfFile
    • FlushViewOfFile可把映射地址空间的内容写到物理文件中
    • UnmapViewOfFile, CloseHandle
  • 管道
    • Windows 提供了匿名管道和命名管道两种管道 机制
    • 利用CreatePipe可创建匿名管道,得到两个读写句柄;利用ReadFile和WriteFile可进行匿名管道的读写
    • 命名管道:一个服务器端与一个客户进程间的通信通道;可用于不同机器上进程通信;作为客户方(连接到一个命名管道实例的一方)时,可以是"\\serverName\pipe\pipename";作为服务器方(创建命名管道的一方)时,只能取 serverName为\\.\pipe\PipeName,不能在其它机器上创建管道
  • 邮件槽mailslot(消息队列):一种不定长、不可靠的单向消息机制,通常采用client-server模式
  • 套接字Winsock: 实现了一个与协议独立的应用编程接口,可支持多种网络通信协议

16.死锁

死锁(Deadlock)是指系统中多个进程无限制地等待永远不会发生的条件

死锁发生的原因: 与不可抢占资源有关

  • 对互斥资源的共享
  • 并发执行的顺序不当

进程使用的资源分为可抢占资源和不可抢占资源两类

  • 可抢占资源(preemptable resource):可以从拥有它的进程中抢占而不会产生任何副作用

的资源。例如 CPU、存储器

  • 不可抢占资源(nonpreemptable resource):在不引起相关的计算失败的情况下,无法把它从占有的进程抢占过来的资源。例如打印机

死锁发生的必要条件

  1. 互斥:任一时刻只允许一个进程使用资源
  2. 请求和保持:进程在请求其余资源时,不主动释放已经占用的资源
  3. 非剥夺:进程已经占用的资源,不会被强制剥夺
  4. 环路等待:环路中的每一条边是进程在请求另一进程已经占有的资源(充分条件)

处理死锁问题的四种方法:

  • 鸵鸟算法:大多数操作系统忽略死锁
  • 死锁预防:预先静态分配法,有序资源使用法
  • 死锁检测:保存资源的请求和分配信息,利用某种算法对这些信息加以检查,以判断是否存在死锁
    • 资源分配图(有向图检测循环)
  • 死锁避免:分配资源时判断
    • 银行家算法(书p258):核心是在试探性分配之前进行安全性检查
* 允许互斥、部分分配和不可抢占,可提高资源利用率;
* 要求事先说明最大资源要求,在现实中很困难

17.处理机调度

处理机调度要解决的问题

  • 按什么原则分配CPU——进程调度算法
  • 何时分配CPU——进程调度的时机
  • 如何分配CPU——进程的上下文切换

调度的开销

  • 从一个进程切换到另一个进程需一定的时间
  • 上下文切换之后,指令和数据高速缓存通常需要更新,执行速度降低 (缺失损失)

非抢先式/抢先式调度算法(时间片+中断)

调度的层次:作业、swap、进程线程

调度算法的目标:

  • 批处理系统
    • 吞吐量——每小时最大作业数
    • 周转时间——从提交到终止的最小时间  CPU利用率——保持CPU忙碌
  • 交互式系统
    • 响应时间——发出命令到得到响应之间的时间(快速响应请求)
    • 均衡性——满足用户的期望
  • 实时系统
    • 满足截止时间——避免丢失数据
    • 可预测性——在多媒体系统中避免品质降低

18.批处理系统中的调度

19.交互式系统中的调度

20.实时系统中的调度

2.3.8 消息传递 消息传递接口MPI “会合”的概念 屏障barrier:

e.g. 大型矩阵运算的分治

避免锁:读-复制-更新 RCU

2.4 调度

处理机调度

I/O活动的含义:阻塞 p85 I/O密集型~多道程序设计 批处理/交互式/实时

系统需要“平衡” p87 公平、平衡、策略强制执行

批处理系统的调度

FCFS
最短作业优先
最短剩余时间优先

轮转调度~上下文切换 https://baike.baidu.com/item/%E4%B8%8A%E4%B8%8B%E6%96%87%E5%88%87%E6%8D%A2/4842616

时间片长度:进程切换用时~响应时间,通常设为20-50ms

关键:衡量指标,用户公平/进程公平/随机/最短进程/CTSS

实时系统的可调度条件 p92

011

速率单调调度RMS 最早截止时限优先EDF 最小裕度算法 laxity

将调度策略(用户态)和调度机制(内核态)分离:参数化

windows:实时优先级线程、可变优先级线程 https://blog.csdn.net/zhiquan/article/details/4240400

线程时间配额 ppt p58

Linux将线程分为三类:

  • 实时FIFO
  • 实时轮转
  • 分时

runqueue 自旋锁 查找O(n)

O(1)调度 亲和性 多核->负载均衡 楼梯调度 抛弃了动态优先级的概念,完全公平 RSDL调度算法 旋转楼梯最终时限调度 (The Rotating Staircase Deadline Schedule) CFS Completely Fair Schedule(完全公平调度)

超级计算机 -- Yunfei Du

应用:

  • 分子动力学:求解牛顿运动方程的积分算法
  • 量子化学计算:用波函数计算薛定谔方程
  • 计算流体力学:Navier-Stokes(NS)方程
  • 计算电磁学:麦克斯韦方程(Maxwell's equation)

体系结构特征:

  • 高性能计算物理结点
    • 单机/双机,不使用虚拟化
    • 并行:向量/多核/大规模异构
    • 体系结构层面:处理并行度、减少数据移动、降低数据精度(比如 TPU bf16,尾数表示和计算机功耗密切相关)
  • 全局共享的并行文件系统
    • Burst Buffer
* SSD实现,在 IO Nodes 和 Compute Nodes 之间
* 处理对IO突发访问需求:比如 checkpoint
* relaxed posix semantics, 和资源管理系统对接
* 要求:一半 memory 的检查点要在 5min 内做完
  • Parallel File System (Lustre 为代表)
* 全局单个命名空间,符合POSIX标准
* MDS/MDT管理和存储所有的元数据
* 多个OSS/OST以对象的方式分布式存储数据
* 分布式锁管理机制:优化了POSIX语义,以优化并行IO
* 元数据和文件数据的通信链路分开管理
  • HPSS (High Performance Storage System)
* 冷数据长期存储
  • 高速互连网络 RDMA
RDMATCP
传输方式消息数据流
延迟0.6us5-50us
带宽200Gb/s40Gb/s
流控Credit based机制滑动窗口机制

IO软件栈

  • 实现 IO 方法
    • Scientific IOlib
    • MPI IO
    • POSIX
  • IO 模式
    • N进程 --- 1个共享文件
    • N进程 --- M个文件

典型超算系统 —— Fugaku

  • 基于arm V8处理器的同构系统
    • 48 cores + 4 assisted cores(4 CMGs)
    • 2*512bit SIMD
    • HBM2 Memory, 1024GB/s bandwidth
    • SoC设计:集成互连网络Tofu
    • 7nm 台积电工艺,表面水冷散热
    • 目前最强性能arm V8处理器,融合AI计算操作
* 3.37Tflops DP(2.2GHz)
  • Fugaku互连网络TofuD
    • Port 带宽:6.8GB/s * 6 RDMA engines = 40.8GB/s
    • Network Interface on Chip
  • Fugaku存储系统
    • 三级层次存储
* GFS Cache + Temp FS(25~30PB NVMe)
* Lustre-based GFS
* Off-site Cloud Storage
  • Fugaku系统组成: 158976结点,414机柜

典型超算系统 —— Summit

  • 2 Power9 CPU <-------- NVLINK --------> 6 V100 GPU,CPU+GPU异构融合

HPC & AI

  • hpc
    • CPU计算为中心
    • 单进程控制一个GPU
  • ai
    • GPU计算为中心
    • 单进程+多个GPU
  • hpc->ai
    • 应用:使用HPC处理训练数据;加速并行训练;推理加速
    • e.g. Fugaku
  • ai->hpc
    • 应用:用于迭代方法挑战收敛参数;减少多目标优化结果的参数空间;当物理模型不清楚或者计算量很大时,AI替代仿真
    • 希望GPU尽可能独立工作,如 gpu direct storage 的设计
    • tpu做科学计算
    • 混合精度解决hpc应用

两种发展思路

  • 共享内存 ~ OpenMP runTime:高性能系统工程量大

<->

分布式系统 内存瓶颈, scaling不好(跨节点延迟高,微秒级) -> nvme ssd


Computer Architecture

[toc]

Intro

性能优化理论

- 业务流程分析:是否有冗余计算,是否能实现场景需求,比如pointwise、listwise - 整体代码性能分析:design是否有更优解,比如pipeline的编排能否更合理、no padding等能否达成 - 关键处代码瓶颈分析:例如folding、unique、gradient reduce等操作

Roofline Model

  • Roofline: an insightful visual performance model for multicore architectures.
    • 横轴:arithmetic intensity: which is the number of arithmetic operations per byte of memory access. (FLOPS/BYTE)
* Compute-bound: the time taken by the operation is determined by how many arithmetic operations there are, while time accessing HBM is much smaller. Typical examples are matrix multiply with large inner dimension, and convolution with large number of channels.
* Memory-bound: the time taken by the operation is determined by the number of memory accesses, while time spent in computation is much smaller. Examples include most other operations: elementwise (e.g., activation, dropout), and reduction (e.g., sum, softmax, batch norm, layer norm).

* 算术操纵数量/内存访问字节数量
* e.g. 
  * 标准attention:O(d)
  * Flash-attn:O(N)
  • 纵轴:computational throughput

image-20250404195013648

  • 多种 workload
    • compute bound
  • memory bound
  • latency bound
  • 硬件例子:
    • V100: 125/0.9 =139FLOPS/Byte
  • Memory optimization和runtime optimization
    • 往往是互相制约的
  • IO-Aware Runtime Optimization 【flash-attention】
    • We draw the most direct connection to the literature of analyzing I/O complexity in this work [1], but concepts of memory hierarchies are fundamental and has appeared in many forms, from the working set model [21], to data locality [86], to the Roofline model of arithmetic intensity [85], to analyses of scalability [59], to standard textbook treatments of computer architecture [40].

并行计算理论

CME323 https://stanford.edu/~rezab/dao/notes/lecture01/cme323_lec1.pdf
  • 分析方法:
  • Assess, Parallelize, Optimize, Deploy(APOD) design cycle
  • 并行算法的“成本”(Cost)通常被认为是 处理器数量 × 时间复杂度 。 在GPU编程中,我们通常分配线程(threads)而不是直接控制物理处理器(processors)。
  • 初始想法 :对于一个包含N个元素的规约操作,一个直接的想法是为每个元素分配一个线程,即使用 O(N) 个线程。
  • 时间复杂度 :使用树形结构进行两两相加的规约,其时间复杂度是 O(log N)。
  • 成本计算 :因此,这种策略的成本是 O(N) * O(log N) = O(N log N)。
  • 结论 :这个成本高于顺序算法的成本 O(N),因此被认为是“非成本高效的”(not cost efficient)。这意味着并行化并没有带来理论上最佳的加速比。
  • Work efficiency的概念
  • A parallel algorithm is work-efficient if it performs the same amount of work as the corresponding sequential algorithm
  • 用于分析scan优化
  • If resources are limited, parallel algorithm will be slow because of low work efficiency.
  • thread coarsening是均衡work efficiency和parallelism的一种方法
* insight:如果parallelism带来了计算serialization的代价,与其硬件支付,不如自己支付serialization
* 适用于:work balance的场景 + memory bound的场景
  • Brent's theorem
  • image-20250509140316594
  • 用于分析reduce sum的优化

Case Study: Unique Kernel Optimization

Reference: PyTorch 工程实践(二):Unique 的性能优化
  • 问题背景:PyTorch CPU 版 unique 算子性能远慢于 NumPy(慢约 6-20 倍)。
  • Architecture Analysis:
*   **CPI (Cycles Per Instruction)**: Profiling 显示 CPI Rate > 5(很差),说明 CPU 流水线停顿严重,通常由**随机访存 (Random Access)** 导致的 Cache Miss 引起。
*   **Bottleneck**: 热点集中在 `std::unordered_set`(哈希表)操作。
  • Optimization Strategy (Sort-based vs Hash-based):
*   **Hash-based (Original)**: 随机内存访问,破坏局部性 (Locality);串行执行,无法利用多核 (Multi-core) 和 SIMD。
*   **Sort-based (Optimized)**: 
    1.  **Sort**: 将相同元素聚在一起,利用连续内存访问。
    2.  **Unique Mask**: `mask[i] = (input[i] != input[i-1])`,无依赖,可完全并行 (Embarrassingly Parallel)。
    3.  **Scan (Prefix Sum)**: 计算输出位置。这是并行计算中的经典原语 (Parallel Primitive)。
    4.  **Scatter**: 并行写入。
  • 收益:
*   **Memory Efficiency**: 避免了细粒度内存分配,利用了 Spatial Locality。
*   **Parallelism**: 全流程可并行,利用多核带宽。

[OSDI 25] 八种系统性能优化方法论

https://www.usenix.org/system/files/osdi25-park-sujin.pdf 《Principles and Methodologies for Serial Performance Optimization》

image-20251019030956778

  • 三大优化原则
原则标识原则名称核心逻辑效果
$$P_{rm}$$任务移除从序列$$S_n$$中删除不必要任务优化后序列长度$$m < n$$
$$P_{rep}$$任务替换用更快任务$$t_j$$替换原任务$$t_i$$序列长度不变,$$F(S_n') < F(S_n)$$
$$P_{ord}$$任务重排序调整任务执行顺序利用 locality 等提升效率

为什么优化 Latency 困难

* Ethernet connection: 0.3ms
* modem link: 100ms
* modem link传输:典型值 33kbit/sec
  • 磁盘的seek time
  • 在低带宽的专用连接和高带宽连接的一小部分份额之间做选择,应该选择后者。

Low Latency Guide

https://rigtorp.se/low-latency-guide/

  • Disable hyper-threading
    • Using the CPU hot-plugging functionality to disable one of a pair of sibling threads. Use lscpu --extended or cat /sys/devices/system/cpu/cpu*/topology/thread_siblings_list to determine which “CPUs” are sibling threads.
    • 由于超线程的存在,CPU 使用率往往是低估的。所以你 CPU 使用率算出来是 50% 时,很可能实际对物理真实核心的占用其实是 70%。
  • CAS Latency: https://en.wikipedia.org/wiki/CAS_latency
    • One byte of memory (from each chip; 64 bits total from the whole DIMM) is accessed by supplying a 3-bit bank number, a 14-bit row address, and a 13-bit column address.
  • 绑核
    • sched_getaffinity、sched_setaffinity
  • Uncore

Arch with C++

constexpr size_t CACHE_LINE_SIZE =
#if __cplusplus >= 201703L and __cpp_lib_hardware_interference_size
  std::hardware_constructive_interference_size;
#else
  64;
#endif

__buildin_prefetch: https://www.daemon-systems.org/man/__builtin_prefetch.3.html

Using the extra 16 bits in 64-bit pointers

CPU

Intro

  • 核心性能因素
    • 单核性能
* 频率 clock speed
  * boost clock
* 流水线设计
* 指令发射数 issue width
  • 核心数量
  • 多级缓存

Branch Predictor

Zen5的2-Ahead Branch Predictor

各种微架构

Intel

Xeon Gold 5118 - Intel

14 nm lithography process

AMD

  • The basic unit of a Ryzen processor is a CCX or Core Complex, a quad-core/octa-core CPU model with a shared L3 cache.
  • However, while CCXs are the basic unit of silicon dabbed, at an architectural level, a CCD or Core Chiplet Die is your lowest level of abstraction. A CCD consists of two CCXs paired together using the Infinity Fabric Interconnect. All Ryzen parts, even quad-core parts, ship with at least one CCD. They just have a differing number of cores disabled per CCX.
  • zen-3(milan)舍弃了一个CCD包含两个CCX的概念,8 cores (one CCD/CCX) 共享 32MB 的 L3 cache
  • Intel’s Monolithic Design and the Future
  • 1 socket ~ 4 CCX ~ 8 memory controller(~ memory channel)
* Up to 2 DIMMs per channel
  • With this architecture, all cores on a single CCD are closest to 2 memory channels. The rest of the memory channels are across the IO die, at differing distances from these cores. Memory interleaving allows a CPU to efficiently spread memory accesses across multiple DIMMs. This allows more memory accesses to execute without waiting for one to complete, maximizing performance.
  • Nodes Per Socket (NPS)
  • For additional tuning details, please refer to the Tuning Guides shared by AMD [here](For additional tuning details, please refer to the Tuning Guides shared by AMD here. For detailed discussions around the AMD memory architecture, and memory configurations, please refer to the Balanced Memory Whitepaper). For detailed discussions around the AMD memory architecture, and memory configurations, please refer to the Balanced Memory Whitepaper

AMD 课程

ROCm:https://developer.amd.com/resources/rocm-learning-center/

内存

DRAM vs SRAM

  • DRAM: 1 transistor, 1 capacitor
  • SRAM: 6 transistors
    • So SRAM is faster but more expensive, takes up more space and gets hotter

image-20250502153520363

Cache 系列科普 ~ Latency

CPU Cache

99f26bb4-bb61-4f00-abf5-99f6a3a82b22

UEFI和BIOS探秘 —— Zhihu Column

interactive latency numbers

数据基于 Skylake 架构

  • L1 cache
    • 32KB
    • 4~5 cycles, L1D latency ~1ns
    • 算一下 load/store 指令占所有 instructions 的比例,小于 5 就没办法 hide latency,需要优化访存模式
  • L2 cache
    • 512KB
    • ~12 cycles, ~4ns
  • LLC (L3 cache) 很关键
    • 32MB
    • ~38 cycles, ~12ns
* 与之对比,memory access 约 50~100ns (Intel 70ns, AMD 80ns)
  • 直接走 IMC,1.5MB/core
  • 分析:llc-load-miss * 64B per load / time elapsed,和内存带宽数据做对比
* 晶体管数目增长落后于晶体管密度增长
* Coffeelake 8700K,晶体管的密度不增反降,Pitch从70nm增加到了84nm。在可以提供更高频率支持的背后,代价就是对Die的大小造成负面影响

Cache Coherence

* 为什么偷显存性能低的原因:显存不能保证被cache,或者说无法保证cache的一致性
  • Cache 一致性
* 用硬件而非软件来做 cache coherency
* CPU 片内总线架构演进:ring bus -> mesh
* 模型:MESI protocol

img

img

SMP (Symmetric Multiprocessing): cache的发展

Bus Snooping (1983)

  • 实现:Home Agent (HA),在内存控制器端;Cache Agent (CA),在L3 Cache端

img

  • 缺点:在QPI总线上广播,带宽消耗大,scaling 有问题
  • Write-invalidate
  • Write-update

The two most common mechanisms of ensuring coherency are snooping and directory-based, each having their own benefits and drawbacks.

  • Snooping based protocols tend to be faster, if enough bandwidth is available, since all transactions are a request/response seen by all processors. The drawback is that snooping isn't scalable. Every request must be broadcast to all nodes in a system, meaning that as the system gets larger, the size of the (logical or physical) bus and the bandwidth it provides must grow.
  • Directories, on the other hand, tend to have longer latencies (with a 3 hop request/forward/respond) but use much less bandwidth since messages are point to point and not broadcast. For this reason, many of the larger systems (>64 processors) use this type of cache coherence.
  • scalability of multi-thread applications can be limited by synchronization
    • 延伸:PCIe 内部 memory (包括 PCIe 后面的显存、NvRAM 等)的割裂性在服务器领域造成了很大问题,CXL 的引入为解决这个问题提供了技术手段
  • synchronization primitives: LOCK PREFIX、XCHG

内存拓扑的 unbalanced 问题

  • 可能导致同一物理机上先启动的服务效率高
  • 多 channel 的 64bit DRAM,ddr 频率在 2666 居多,单 channel 可以到 ~20GB/s,4~6 channel 比较常见

mem-layout

Some conclusion and Advices

  • 11 -> 1H: Hyper Threading could help on performance on such “lock” condition (But may not in the end, maybe depends on the total threads: C1-> CH)
  • 22 -> 21: Lower Core-Count Topology helps for this circustances (Not the Benchmark Software threads)
  • Increase of the hardware resource (along with a huge mount of OS threads) usually not help on the performance, but waste of the CPU and memory resources
  • Intel’s “innovation” for high performance processors has been tired with maintaining the same performance for “unconstrained” end users.
    • Intel has been done very well, if you compare with ARM64 Enterprise and AMD...
  • lock code is the RISK (pitfall from something beyond your source code, even from glibc, 3rd lib ...)
  • Use the scaling tests to find your bottleneck, and improve the “lock” components
    • Maybe from DISK I/O, Network layer
    • Rarely from the memory bandwidth layer, LLC cache size for the non-HPC workloads

TLB

  • 内存页条目缓存即TLB( Translation Lookaside Buffer)TLB缓存命中率越高,CPU执行指令的速度越快。
  • cat /proc/interrupts | grep "TLB shootdowns"
    • 文件记录的是 remote TLB flush 事件。本地CPU使用IPI中断通知其他CPU flush TLB时,该节点对应的CPU会计数。
  • perf stat -e cache-misses,cache-references,instructions,cycles,faults,branch-instructions,branch-misses,L1-dcache-stores,L1-dcache-store-misses,L1-dcache-loads,L1-dcache-load-misses,LLC-loads,LLC-load-misses,LLC-stores,LLC-store-misses,dTLB-loads,dTLB-load-misses,iTLB-loads,iTLB-load-misses -p $PID

NUMA

  • 搜到一段获取 cpu topology 的 C 语言代码:https://github.com/SANL-2015/SANL-2015/blob/8779af7939bcacebd74abfabba9873b68eaca304/SAND2015/liblock/liblock.c#L99
yum -y install numactl numastat

numactl -H
numastat
numactl -C 0-15 ./bin
numactl -N0 -m0 ./bin

Hyper-threading

https://www.semanticscholar.org/paper/Hyper-Threading-Technology-Architecture-and-1-and-Marr-Binns/04b58af4fc0e5c3e8e614e2ddb0c41749cc9166c

https://pdfs.semanticscholar.org/04b5/8af4fc0e5c3e8e614e2ddb0c41749cc9166c.pdf?_ga=2.24705338.1691629142.1553869518-295966427.1553869518

https://www.slideshare.net/am_sharifian/intel-hyper-threading-technology/1

  • 实测性能是 -20% ~ +20%,因为可能依赖内存带宽、抢占cache。数值计算任务,用到了AVX、SSE技术的,开Hyper-threading一般都会降性能;访存频繁的CPU任务,建议打开,因为circle、LRU是idle的,会有提升
  • 禁止hyper-threading:offline、isolate

指令集

汇编优化

  • FFmpeg:https://github.com/FFmpeg/asm-lessons

AMX

容量成本低,带宽成本高

image-20251005230828442

  • AMX: The x86 Advanced Matrix Extension (AMX) Brings Matrix Operations; To Debut with Sapphire Rapids
    • AMX introduces a new matrix register file with eight rank-2 tensor (matrix) registers called “tiles”.
    • 独立单元,支持bf16,intel oneAPI DNNL指令集接口
  • 功能:单元相比avx512、vnni,支持了reduce操作
  • 实现:
    • 输入8位/16位,计算用32位(防溢出)
    • 扩展了tile config、tile data寄存器,内核支持(XFD,eXtended Feature Disable), allos os to add states to thread on demand
  • 测试:减少内存带宽,增频了(重计算指令减少)
    • 配合tf有automix策略,op可能为bf16+fp32混合计算
    • matmul相关运算全bf16
    • tf2.9合入intel大量patch, TF_ENABLE_ONEDNN_OPTS=1,检测cpuid自动打开;oneDNN 2.7
- TF_ONEDNN_USE_SYSTEM_ALLOCATOR
- jemalloc: MALLOC_CONF="oversize_threshold:96000000,dirty_decay_ms:30000,muzzy_decay_ms:30000"), mitigate page_fault, which benefit fp32 in addition
- TF_ONEDNN_PRIM_CACHE_SIZE=8192 (for mix models deployment)
- TF2.9 SplitV performance issue
  • 其它SPR独立单元:
    • Intel DSA(data streaming accelerator): batched memcpy/memmove,减少CPU cycles
- https://01.org/blogs/2019/introducing-intel-data-streaming-accelerator
  • Intel IAA(in-memory analytics accelerator): compress/decompress/scan/filter,也是offload cpu cores
- https://www.intel.com/content/www/us/en/analytics/in-memory-data-and-analytics.html
- 场景如presto:https://engineering.fb.com/2019/06/10/data-infrastructure/aria-presto/
应用场景:Mooncake KTransformers

image-20251005231348062

存储:硬盘、NVMe

  • 读写速度
  • 万兆网(10Gb/s):理论带宽是1.25GB/s, 实际一般能达到800~1200MB/s
  • SATA SSD:读取速度通常在500MB/s 左右
  • NVMe SSD(PCIe 3.0): 读取速度可以达到3500MB/s 甚至更高
  • NVMe SSD(PCIe 4.0): 读取速度可以达到7000MB/s 或更高
  • 硬盘:分为HDD和SSD
    • 读写模式
* 随机读写:高频、小文件
  * IOPS
  * 瓶颈:磁盘读取到缓冲区
* 连续读写:顺序、大文件
  * MB/s
  * 瓶颈:缓冲区传输到内存
  • 典型数据(NVMe):
* 1TB
* 700000 IOPS
* 7000 GB/s
  • 优化硬盘读写的技术:
* 利用内存
  * mmap,尤其随机访问场景
  * RAMDisk
* 尽量连续读写
* NVMe
* 使用传输速率更高的PCIe通道
* https://www.youtube.com/watch?v=MSaD8DFsMAg
  • Persistent Memory
* Optane DIMM: https://www.anandtech.com/show/12828/intel-launches-optane-dimms-up-to-512gb-apache-pass-is-here
  * Optane DC PMMs can be configured in one of these two modes: (1) memory mode and (2) app direct mode. In the former mode, the DRAM DIMMs serve as a hardware-managed cache (i.e., direct mapped write-back L4 cache) for frequently-accessed data residing on slower PMMs. The memory mode enables legacy software to leverage PMMs as a high-capacity volatile main memory device without extensive modifications. However, it does not allow the DBMS to utilize the non-volatility property of PMMs. In the latter mode, the PMMs are directly exposed to the processor and the DBMS directly manages both DRAM and NVM. In this paper, we configure the PMMs in app direct mode to ensure the durability of NVM-resident data.
  * pmem: https://pmem.io/pmdk/
  * DWPD: 衡量 SSD 寿命
  * 《Spitfire: A Three-Tier Buffer Manager for Volatile and Non-Volatile Memory》
  • 内存
    • 双通道、四通道,将内存带宽提升相应倍数

jemalloc

内存分配器

  • 栈内存的生命周期是函数,堆内存可能是进程
  • 系统调用细节
    • C++14 开始支持 sized free,要求 size 和指针对应内存申请时的大小相同
  • size classes
    • 4KB逻辑页 -> "small" 小于 16KB

jeprof使用

export MALLOC_CONF="prof:true,prof_leak:true,lg_prof_interval:31,prof_final:true,prof_prefix:jemalloc/jeheap"

apt-get install -y binutils graphviz ghostscript

jeprof --show_bytes --pdf /usr/bin/python3 [--base=98/heapf_.3408.0.i0.heap] 98/heapf_.3408.44.i44.heap > result.pdf

je 接口与参数

非标准接口
  • mallocx / rallocx 返回已分配的内存指针,null表示没有符合条件的内存
    • realloc(), rallocx, xallocx : in-place resizing
  • xallocx 返回 ptr resized 结果
  • sallocx 返回已经分配的 ptr 的真实大小
  • nallocx 返回 mallocx 可以成功的试算大小
  • mallctl、mallctlnametomib 和 mallctlbymib 控制 jemalloc 内部状态
  • dallocx= free
  • sdallocx = sized free
调参

http://jemalloc.net/jemalloc.3.html#TUNING

  • opt_dss(用法 MALLOC_CONF=dss:primary ): primary 表示主要使用brk,seconday 表示优先mmap,默认是 seconday
内存泄漏分析

https://zhuanlan.zhihu.com/p/138886684

gdb call malloc_stats_print(0, 0, 0)

jemalloc 多篇论文介绍

“Understanding glibc malloc” by Sploitfun. February, 2015.”

  • per thread arena
  • Fast Bin; Unsorted Bin; Small Bin; Large Bin; Top Chunk; Last Remainder Chunk

(jemalloc) A Scalable Concurrent malloc(3) Implementation for FreeBSD

Introduction: allocator 性能逐渐成为重点

问题1: 要尽量避免 cache line 的多线程争抢问题

  • jemalloc instead relies on multiple allocation arenas to reduce the problem, and leaves it up to the application writer to pad allocations in order to avoid false cache line sharing in performance-critical code, or in code where one thread allocates objects and hands them off to multiple other threads.
    • Multiple Arenas 默认是cpu core 数乘以 4
    • 线程争抢 (false sharing):一个 cacheline 64 bytes,如果a线程b线程每个线程共享一块malloc出来的内存的前32byte和后32byte,两个线程会争抢这个cache line。这种编程模式,依赖业务代码加一些 padding 避免争抢

问题2: reduce lock contention <-> cache sloshing

  • jemalloc uses multiple arenas, but uses a more reliable mechanism than hashing for assignment of threads to arenas.

解决方案:

  • the arena is chosen in round-robin fashion,用 TLS(Thread-local storage) 实现,可被 TSD 替代
  • 理解 "chunk" 2 MB,small、large、huge,huge 用单个红黑树管理
  • 对于small/large,chunks 用 binary buddy algorithm 分割成 page runs
    • small allocations 用 bitmap 管理,run header 方案的优劣讨论
    • fullness 这段没看懂

Scalable memory allocation using jemalloc by Facebook

生产环境的 challenges for allocations

  • Allocation and deallocation must be fast.
  • The relation between active memory and RAM usage must be consistent.
  • Memory heap profiling is a critical operational aid.

文章介绍了 jemalloc 的核心设计思想:通用地去解决过往 allocator 解决/待解决的问题

img

img

Tick Tock, malloc Needs a Clock

  • Background & basics
    • Chunk: 1024 contiguous pages (4MB), aligned on a 4MB boundary
    • Page run: 1+ contiguous pages in a chunk
    • Region: contiguous bytes (<16KB)
  • Fragmentation & dirty page purging
    • 重点关注 external fragmentation
    • 设计原则:prefer low addresses during reuse (e.g. cache bin), exceptions:
* size-segregated slabs for small allocations (size class 互相独立,不去找更前面的 size class 制造 internal fragmentation)
* disjoint arenas
* thread caches
* unused dirty page caching delays page run coalescing
  • Clockless purging history
    • jemalloc 不进行异步调用,all work must be hooked into deallocation event
    • side effect: conflict between speed and memory usage optimization
    • Aggressive (2006): immediately purge, disabled by default, MADV_FREE-only
    • Spare Arena Chunk (2007):
* keep max one spare per arena, which never purges
* Rational: hysteresis for small/huge allocations
  • Fixed-size dirty page cache (2008): 512 pages per arena,追求速度
  • Proportional dirty page cache (2009): active/dirty=32/1,内存更大
  • Purge chunks in FIFO order (2010)
* Chunk iterations cost a lot
* Rational: purge chunks with high dirty page density (faster)
  • Purge chunks in fragmentation order (2012)
* maintain chunk fragmentation metric
  • Purge runs in LRU order (2014)
* Dirty runs 进入 LRU 后 coalesce, 减少了整体的 madvise 次数
  • Integrate chunks into dirty LRU (2015)
* Rational: huge allocation hysteresis
  • Clock-based purging challenges
    • Prototype
* ~4.3s delay (2e32 ns)
* 32-slot timer wheel
  • 结果内存涨的太多,因为尤其是多个 arena 内存的累加
  • Limit dirty page globally?
* too strict -> per numa node attached RAM
* Jemalloc 目前没有将 CPU core 和 arena 联系起来
  • Hybrid? 不现实
  • decay curves
* timer wheel stutters
* Sigmoid > delayed linear > linear >> exponential
  • Online decay algorithm
* Asynchronus purging thread
* Synchronus purging during allocation at peak usage
  • Prefer dirty run reuse? 期望不要,因为会 purge 更频繁
  • Other wall clock uses
* Incrementally flush arena cache
* Incrementally/asynchronusly flush thread cache
  * Idle threads don't need caches
  * Current HHVM hack: LIFO worker thread scheduling; flush cache after 5 seconds inactivity
  * Flushing may be impossibly without impacting fast path (必须和线程做同步)

jemalloc 代码阅读

img

  • malloc(), posix_memalign(), calloc(), realloc(), and free()
  • rudimentary introspection capabilities: malloc_usable_size()
  • *allocx(): Mix/match alignment, zeroing, relocations, arenas, tcaches
  • mallctl*(): Comprehensive introspection/control
  • base_alloc(): metadata,jemalloc 直接调用 mmap 从 os 申请,不释放
  • malloc(): 递进的 size class(默认页 4KB)
  • free(): lib 自己维护地址到 metadata 的映射
    • ptmalloc: 在每一个 malloc 出去的内存块前都加若干字节的 metadata (pre size field - chunk size field - N - M - P - User Data - next chunk),所以相邻的内存块的溢出可能会踩踏到下一块内存的 metadata
    • jemalloc: 内存地址稀疏,radix trees
  • jemalloc 设置了一个界线,size < 4*os_page_size 的 size_class, 称这类分配为"small",内存管理会用 bitmap 额外再多一层 slab 缓存。size>=4*os_page_size 的 size_class 称为 "large",不再进行额外的缓存。
  • TLS cache
    • x86 的 linux 下会使用 fs 段寄存器指向线程的 TCB
    • 动态链接库访问线程局部变量理论上要经过 __tls_get_addr 系统调用,有 overhead,而 jemalloc 绕过了这一开销。缺点是 jemalloc 这么编译出来的 lib 是没法被 dlopen 动态链接的,详情见 --disable-initial-exec-tls 这个 flag
    • cache bin: include/jemalloc/internal/cache_bin.h:cache_bin_alloc_easy()
    • maintains large objects up to a limited size (32 KiB by default),增加这个 limit 会导致 unacceptable fragmentation cost
    • 对应地,有 GC 策略:Cached objects that go unused for one or more GC passes are progressively flushed to their respective arenas using an exponential decay approach.
  • bin
    • include/jemalloc/internal/bin.h
    • 开始并发访问,thread 到 arena 的分配是 round robin
    • large size class: 下一层 extent 的包装
* `src/large.c:large_malloc()`
  • small size class: 如果没取到 cache_bin,tcache_alloc_small_hard()->arena_tcache_fill_small()->arena_slab_reg_alloc_batch()
* `bin->slancur` 指向当前可用的 bitmap
* nonfull_slab/full_slab: pairing heap
* jemalloc 希望分配的内存空间尽量紧凑,地址尽可能复用,这样能获得更好的 cache line,TLB 的局部性,所以用到 Pairing heap 维护 slabs_nonfull 去查找尽可能老的,或地址空间最低的 slab 用于分配
* `arena_slab_reg_alloc_batch()`: bitmap 操作, [ffsl](https://man7.org/linux/man-pages/man3/ffs.3.html) find first set
  • extent
    • dirty muzzy retained
* `extents_t extents_dirty`
* large size class 的核心函数:`arena_extent_alloc_large()`
* small size class 的核心函数:`arena_bin_malloc_hard()--->arena_bin_nonfull_slab_get()--->arena_slab_alloc()`
  • slab/large_alloc <---> extents_dirty <---> extents_muzzy <---> extents_retained
  • 内存块的合并:每次从左边的 extent 向右边释放的时候,会查询全局的 radix trees,检查这个 extent 是否能和相邻的 extent 合并
  • 调用 madvise 定时对内存进行 gc, 可能会引起系统的 stall
  • 小内存会抢占大内存,切割大内存的 extent;分配完也会合并成大内存 extent
  • kernal
    • 内存地址空间的分配 mmap
* linux虚地址空间是由task_struct对应的mm_struct指向的vm_struct链表管理的。mmap系统调用在虚地址空间查找一段满足length长度的连续空间后,创建一个vm_struct并插入链表。mmap返回的指针被程序访问时将触发缺页中断,操作系统分配物理页到vm_struct中,物理内存的增加体现在rss上。
* `mmap()<---os_pages_map()<---pages_map()<---extent_alloc_mmap()`
* extents_retained 只存虚拟地址空间,没有物理页
* `pages_map()` 的逻辑主要在处理内存 alignment
  • 内存清理,认为 unmap 有开销,于是只清理物理内存,不清理虚拟地址空间
* 代码详见 `base_unmap()`
* extents_dirty ---> extents_muzzy 调用 `madvise(MADV_FREE)`,轻量操作,将内存给其它内存
  * `pages_can_purge_lazy`
* extents_muzzy ---> extents_retained 调用 `madvise(MADV_DONTNEED)` ,将内存拿走,可能发生ipi中断,再次访问会产生缺页中断
  * `pages_can_purge_forced`
* Facebook 优化:`mmap(...MAP_UNINITIALIZED)`
  • 透明巨页的分配:从mmap出来的extern,pages_huge_impl()调用 madvise(MADV_HUGEPAGE)
  • opt_dss支持brk方式申请
  • 其它
  • 性能瓶颈:cache bin 的 fill、flush,madvise;arena 分配机制(场景:多个线程分配一个静态对象内的内存,单线程操作它,会产生 high fragmentation)
  • jemalloc_internal_defs.h.in: 一串 undef + 宏的含义注释
高性能内存分配库 Libhaisqlmalloc 的设计思路

https://zhuanlan.zhihu.com/p/352938740

显示器

一些外设概念

  • OSD(On Screen Display) Menu
  • 接口:DC-IN, HDMI 2.0 两个, DisplayPort, 耳机, USB 3.0 两个, Kensington 锁槽
  • 170Hz 刷新率、130%sRGB、96%DCI-P3

显卡

gpu-z 判断锁算力版本

主板

  • PCI-E x1/x4/x8/x16
    • PCI-E x16:22(供电)+142(数据);用于显卡,最靠近 CPU
    • PCI-E x8:伪装成 x16
    • PCI-E x4:22+14;通常由主板芯片扩展而来,也有直连 CPU 的,用于安装 PCI-E SSD
    • PCI-E x1:独立网卡、独立声卡、USB 3.0/3.1扩展卡等
* 另外一个形态,一般称为Mini PCI-E插槽,常见于 Mini-ITX 主板以及笔记本电脑上,多数用来扩展无线网卡,但由于其在物理结构上与 mSATA 插槽相同,因此也有不少主板会通过跳线或者 BIOS 设定让 Mini PCI-E 接口在 PCI-E 模式或者 SATA 模式中切换,以实现一口两用的效果。已经被 M.2 接口取代,基本上已经告别主流。

通信与网络

通信与网络

lecture 8.线性分组码译码及应用

CRC循环冗余校验码 (cyclic redundancy checks)

  • (k+m,k)码,$r(x)=x^{m}d(x)\mod g(x)$, $c(x)=x^md(x) - r(x)$
  • 当g(x)的重量不为1时,CRC码距大于等于2
  • 存在周期: $ x^t-1 \mod g(x) = 0$
    • m+k>n时:存在两列相同 => CRC无纠错能力,只能检错 ; 最小码距为2
    • m+k=n:循环码, (10, 7)码不能纠错, (7, 4) 码可以; e.g. g(x)=1101 周期为7
  • CRC is linear,密码学的内容=>没有integrity
  • 可检查突发错(burst): 检错能力比下限强

5G时代的机遇与挑战 —— 吕廷杰

回顾历史

  • 《第三次浪潮》声明其为新的时代,智力密集型行业
  • 互联网
    • 美苏核对峙 ---> 除核之外,通信是军事中最重要且脆弱的部分(层级式通信,单点 fail)
    • p2p 网络,1969年互联网,研究了 packet switching 分组交换技术、sliding window ---> 速递包裹的跟踪查询系统
    • PC 互联网 ---> 移动互联网
  • Q: 互联网是否创造了价值?
    • A: 移动互联网拥抱实体经济,重新定义了生产力
  • 1G->2G: 能力完善
  • 2G->3G: Apple 重新定义了移动互联网,App Store 是 OTT (over-the-top) service,网络和业务分离
  • 3G->4G: 降价提速

5G

时间定义交换带宽
1G1980语音通信无线化
2G1990无线通信数字化100Kbps
3G2000无线通信上网化100Mbps
4G2010上网手机娱乐化1Gbps
5G2020万物互联10Gbps
  • 5G 的风险不在于技术层面,而在于应用层有多大的创新空间,尤其是与国家战略相关的创新
  • 5G 的三大应用
  • eMBB (enhanced Mobile Broadband)
* 数字化娱乐、UHD screen、远程手术、智能家居办公
  • URLLC (Ultra Reliable Low Latency Communications)
* 工业自动化、AR/VR、
  • mMTC (massive Machine Type Communications)
* 自动驾驶、智能机器人、无人机、智能制造
  • 5G 的四大商业模式
  • 基于流量:eMBB, 2C;需要运营商加快用户分级的智能管道升级,实现差异化的流量收费模式
  • 基于连接:单独提供连接,可能包括终端设备和模组;运营商按物联网设备采用卡用户收益(月/年)等方式收费
  • 基于网络切片:运营商根据不同垂直行业和特定区域定制网络切片,垂直用户可直接向运营商购买网络切片,一般按年计费
  • 基于完整解决方案:为工业企业提供包括工厂内外连接、设备终端数字化改造、平台层完整解决方案等
* ToC 难以提供通用解决方案(难以做好企业画像),给了中小企业杀出来做 ToB 的机会
* 面向行业细分市场的服务不是运营商的强项,因此可以考虑 ToC 转 ToB 能力开放、鼓励分销与转售
  • 5G 将推动边缘计算的发展与应用
  • 见下文 边缘计算 小节
  • 5G 将重构数字经济的产业生态
  • 万物互联,做数据的更挣钱,因为有更多数据
  • 5G 应用与网络部署的进度、垂直行业解决方案、国家政策推动密不可分
  • 5G 与区块链
  • 区块链作为 5G 的账本
  • 5G 的隐私泄漏问题:《维基经济学》,以区块链为终极解决方案;工业企业不用运营商的频率,专用网络用特定频率
  • 5G + AI 消灭工作机会? ---> 多能力组合的文化创意产业

边缘计算

  • 边缘计算
    • 从集中到分布、从被动到主动、减少网络时延、本地化服务
    • 靠近客户、靠近场景、资源独占
    • 一篇综述,分类:
* 物联网边缘计算:边缘计算是物联网的附属品
* P2P边缘计算:“特定内容借流量”和“特型应用借算力”,轻量、不稳定
* 服务器边缘计算:边缘网络 + 边缘IaaS算力 + 网站服务,这就是 CDN;把网站服务换成视频服务,这就是点直播;如果把应用层换成通用边缘计算框架,再通过 5G 把延时降低到 10ms 以下,这就是边缘计算
* 运营商边缘计算
  • 客户价值
* 硬件设计更灵活
* 改变应用发布生态
* 单一应用留住客户
  • 边缘计算场景
    • 视频优化加速:边缘部署给中心视频服务器提供动态网络分析信息,辅助TCP拥塞控制、码率匹配,改善内容分发效率低下情况
    • 监控视频流分析:在边缘进行监控视频的分析,降低视频采集设备的成本,减少发给核心网的流量
    • 车联网:MEC应用分析车及路侧传感器的数据,预警周边车辆,为用户提供路障通知、减小拥堵、感知其他车辆行为等服务
    • 企业专网:在网络边缘对用户接入进入控制和分流,用户面流量分流到企业网络的服务
    • IOT/工业互联:MEC应用整合、分析设备产生的消息及时进行决策,对终端的远端接入和控制等。例如在机场,飞机发动机传感器在着陆前检测状态,AI眼镜探测问题
    • 辅助时延敏感计算:字节跳动 Pitaya,推荐/广告端上智能场景
* “算法包”扫二维码调试、动态部署:算法开发与客户端开发解耦
* 端上特征工程:特征实时性强、可回传云端形成闭环
* 算法包的运行触发:Applog Event 或 业务逻辑主动触发
* 通用 AI 能力建设:针对通用性的使用场景(网络状态预测等),可内置相关能力,快速推广至业务方
  • AR/VR:边缘应用快速处理用户位置和摄像头图像,实时提供辅助信息,本地化处理增强体验

Computer Networking Lab CS144 Stanford

[toc]

CS144-Lab-Computer-Networking

写在前面

在历史的伟力面前,个人的命运是不可捉摸的。学生生涯结束地比想象中快,下个月就要正式入职字节跳动了。回顾本科期间,做过的大作业不少,却大多是期中对着一页不明就里的薄纸发呆,期末临近deadline,东抄抄西补补,勉强弄个不忍直视的半成品,没有时间也没有能力完成一次高质量的大作业。打算在入职前至少做一个Lab,之所以选择stanford的CS144,一方面是因为这门课质量很高,b站有配套的视频,Lab也在这两年做了大的改进,改为了一个优雅的TCP实现,全部资料都开源在课程网站上,适合自学。另一方面,我本科期间没有学过计算机网络,这次补课也有一举两得的意味。

做下来感觉不错,课程的老师和助教很用心:说明文档十页长,FAQ覆盖了作业中会遇到的方方面面的问题;大作业的单元测试的代码量是模块代码的几倍,倾注了助教的心血。如果没有这些模块架构和单元测试,对于初学者来说是很难完成即使是初级的TCP协议栈编写,充分利用这些资源,站在巨人的肩膀上,能力能得到更好的锻炼。

以下是一些课程资源:

  • ☑ Lab 0: networking warmup
  • ☑ Lab 1: stitching substrings into a byte stream
  • ☑ Lab 2: the TCP receiver
  • ☑ Lab 3: the TCP sender
  • ☑ Lab 4: the TCP connection
  • ☐ Lab 5: the network interface
  • ☐ Lab 6: the IP router

Lab0: networking warmup

1.配环境

设虚拟机,实验指导书 ,可参考我的Shell笔记,迁移dotfiles

sudo apt-get update
...
cd sponge/build
rm CMakeCache.txt
CLANG_TIDY=clang-tidy-6.0 CXX=clang++-6.0 cmake .. -DCMAKE_BUILD_TYPE=Debug
2.Networking by Hand

2.1 Fetch a Web page

telnet cs144.keithw.org http
GET /hello HTTP/1.1 # path part,第三个slash后面的部分
Host: cs144.keithw.org # host part,`https://`和第三个slash之间的部分
  • 返回的有ETag, 减少服务器带宽压力
HTTP/1.1 200 OK
Date: Sat, 23 May 2020 12:00:46 GMT
Server: Apache
X-You-Said-Your-SunetID-Was: huangrt01
X-Your-Code-Is: 582393
Content-length: 113
Vary: Accept-Encoding
Content-Type: text/plain

Hello! You told us that your SUNet ID was "huangrt01". Please see the HTTP headers (above) for your secret code.
netcat -v -l -p 9090
telnet localhost 9090
3.Writing a network program using an OS stream socket

OS stream socket: ability to create areliable bidirectional in-order byte stream between two programs

  • turn “best-effort datagrams” (the abstraction the Internet provides) into“reliable byte streams” (the abstraction that applications usually want)

3.1 Build

3.2 Modern C++: mostly safe but still fast and low-level

  • 读文档:https://cs144.github.io/doc/lab0/inherits.html
    • a Socket is a type of FileDescriptor, and a TCPSocket is a type of Socket.
//! \name
//! An FDWrapper cannot be copied or moved
//!@{
FDWrapper(const FDWrapper &other) = delete;
FDWrapper &operator=(const FDWrapper &other) = delete;
FDWrapper(FDWrapper &&other) = delete;
FDWrapper &operator=(FDWrapper &&other) = delete;
//!@}

3.4 webget()

  • SHUT_RD/WR/RDWR,先用SHUT_WR关闭写,避免服务器等待
void get_URL(const string &host, const string &path) {
    TCPSocket sock{};
    sock.connect(Address(host,"http"));
    string input("GET "+path+" HTTP/1.1\r\nHost: "+host+"\r\n\r\n");
    sock.write(input);
    // cout<<input;
    // If you don’t shut down your outgoing byte stream,
    // the server will wait around for a while for you to send
    // additional requests and won’t end its outgoing byte stream either.
    sock.shutdown(SHUT_WR);
    while(!sock.eof())
        cout<<sock.read();  
    sock.close();
}

3.5 An in-memory reliable byte stream

  • 数据结构deque,注意eof的判断条件即可

lab1: stitching substrings into a byte stream

3.Putting substrings in sequence
  • assemble数据时,为了简化代码流程,先将其和可能的字段合并,再判断是否可以write,因此需要设计一个merge函数
  • 注意end_input()的判断条件
  • 用set保存(index, data)数据,可以用lower_bound查找
  • 细节:push_substring的bytes接收范围图
reassembler

lab2: the TCP receiver

3.1 Sequence Numbers
different index
  • 利用头文件中的函数简化代码
  • 计算出相对checkpoint的偏移量之后,再转化成离checkpoint最近的点,如果加多了就左移,注意返回值太小无法左移的情形。
uint64_t unwrap(WrappingInt32 n, WrappingInt32 isn, uint64_t checkpoint) {
    uint32_t offset = n - wrap(checkpoint, isn);
    uint64_t ret = checkpoint + offset;
    // 取距离checkpoint最近的值,因此判断的情况是否左移ret
    //注意位置不够左移的情形!!!
    if (offset >= (1u << 31) && ret >= UINT32_LEN)
        ret -= (1ul << 32);
    return ret;
}
3.2 window
  • lower: ackno
  • higher~window size
  • window size = capacity - ByteStream.buffer_size()
3.3 TCP receiver的实现
  1. receive segmentsfrom its peer
  2. reassemble the ByteStream using your StreamReassembler, and calculate the
  3. acknowledgment number (ackno)
  4. and the window size.
  • _reassembler忽视SYN,所以要手动对index减1、ackno()加1
  • 非常规路线的处理:比如对于第二个SYN或者FIN信号,接收机选择忽视,具体见bool TCPReceiver::segment_received(const TCPSegment &seg)的实现

lab3: the TCP sender

3.1 重传时机

sponge网络库的设计,TCP的测试中利用到状态判断,但具体到sender、receiver这几个类的设计时,对类的设计是面向对象,对类内部的函数是面向过程,而不像Linux内核的tcp实现中有利用goto语句来模拟有限状态机。因此,在实现这个Lab的函数的时候,依然要以面向过程的思路,理解sender和receiver在不同的情景下会如何工作,使函数在内部看来是一个过程,外界测试时又能完美的体现状态变化。比如三次握手和四次挥手就是一个很好的例子

  • sender发送new segments(包含SYN/FIN),用_segments_out这个queue跟踪,影响它的因素是ackno
  • 重传条件是"outstanding for too long", 受tick影响,tick仅由外部的类调用,sender内部不调用任何时间相关的函数
  • retransmission timeout(RTO),具体实现是RFC6298的简化版
    • 重传连续的之后double ;收到ackno后重置到_initial_RTO
    • 可参考RFC 6298第5小节实现_timer
  • 注意读/lib_sponge/tcp_helper/tcp_state.cc帮助理解状态变化
receiver sender

lab4: the summit (TCP in full)

这次Lab是把之前的receiver和sender封装成TCPConnection类,用来进行真实世界的通信,下图有助于直观理解结构。

TCP dataflow TCP header

实验的测试文件一如既往的重要,命名规则如下:

  • “c” means your code is the client (peer that sends the first syn)
  • “s” means your code is the server.
  • “u” means it is testing TCP-over-UDP
  • “i” is testing TCP-over-IP(TCP/IP).
  • “n” means it is trying to interoperate with Linux’s TCP implementation
  • “S” means your code is sending data
  • “R” means your code is receiving data
  • “D” means data is being sent in bothdirections
  • lowercase “l” means there is packet loss on the receiving (incoming segment) direction
  • uppercase “L” means there is packet loss on the sending (outgoing segment) direction.

实现时的重要细节:

1.需要单独讨论重传ACK的情形,在我的实现中我写了一个send_ack_back()函数

In TCPConnection::segment_received, what are the three conditions in which the TCPConnection needs to make sure that the segment receives at least one ACK segment in reply, and may need to force the TCPSender to spit out an empty segment to make this happen?

  • If the incoming segment occupies any sequence numbers (length_in_sequence_space() > 0)
  • If the TCPReceiver thinks the segment is unacceptable (TCPReceiver::segment_received() returns false)
  • If the TCPSender thinks the ackno is invalid (TCPSender::ack_received() returns false)

2.处理RST

  • 如果收到RST,需要给sender和receiver的stream用set_error(),不需要回传
  • 发送RST的情形
    • 错误的connect,connect()函数中
    • unclean shutdown,析构函数中
    • 连续重传超次数 _sender.consecutive_retransmissions() > TCPConfig::MAX_RETX_ATTEMPTS

3.判断终结条件

  • 具体实现:见代码以及实验指导书的第5节。
  • 背后的理念:Because of the Two Generals Problem, it’s impossible to guarantee that both peers can achieve a clean shutdown

Debug

1.最后和linux系统真实地进行通信,总有10个tests过不了,打算先研究这一个测试样例: ../txrx.sh -isDnd 128K -w 8K -l 0.1

2.Github上找到一个顺利过关的印度大哥,他给了一些debug建议

3.改了一些tcp_connection和tcp_sender的细节,benchmark速度提升

4.debug无果,决定替换印度大哥模块逐一排查,发现是tcp_receiver出了问题,

4.最终的bug非常坑,即使是助教写的测试样例也无法照顾到,只有当和linux的tcp进行真实丢包通信时才会出现,具体是在以下这一行tcp滑窗控制:

bool inbound = (seq_start>=win_start&& seq_start<=win_end) || (payload_end>=win_start && seq_end<=win_end);

receive的条件是数据片段和receiver的窗有重合,这个实现看似很简单,只需要写出数据的begin和end,窗的begin和end,然后做判断。具体来说,判断数据的begin是否在窗里,再判断数据的end是否在窗里即可。但在TCPReceiver的具体实现中,在判断数据的end是否在窗里时,需要用两个不同的end:代码中的payload_end是原先的seq_end经过处理得到的,需要把syn和fin占的位置排除掉,这样处理是因为receiver内部的_reassembler类只处理data,不处理syn和fin。

综合来说,需要对基础的窗的判断条件做修改,具体实现中不是一个uniform的形式,只有这样才能过最后几个tests。

lab5: down the stack (the network interface)

  • 用处:传输TCP、router
  • 一些传输TCP Segment的方式
    • TCP-in-UDP-in-IP
* kernel API,确保 isolation、exclusive combination of local and remote addresses and port numbers
  • TCP-in-IP
* linux TUN device: application supply an entire Internet datagram, and the kernel takes care of the rest (writing the Ethernet header, and actually sending via the physical Ethernet card, etc.). But now the application has to construct the full IP header itself, not just the payload
  • TCP-in-IP-in-Ethernet
* Network inferfaces: eth0, eth1, wlan0,将IP datagram转成raw Etherne frames传给TAP device
* Example
* ARP probe
* ARP announcements
* A protocol is needed to dynamically distribute the correspondences between a <protocol, address> pair and a 48.bit Ethernet address.
* assumes a reply is only provoked by a request -> 不用检查op_code::reply,直接插入translation table
* The driver consults the Address Resolution module to convert <ET(IP),
IPA(Y)> into a 48.bit Ethernet address, but because X was just
started, it does not have this information.
* 讨论了table aging and/or timeouts的条件:建连失败、收包超时
  * the only bad information that can exist is in a machine that doesn't know that some other machine has changed its 48.bit Ethernet address.  Perhaps manually resetting (or clearing) the address mapping table will suffice.

lab6: building an IP router

lab7: put it all together

工程细节

  • 注意迭代器的使用
    • container.erase(iter++), 同时完成删除和迭代
    • 如果iterator重复erase,可能导致seg fault
  • 单元测试
    • 用generate生成随机数据
auto rd = get_random_generator();
const size_t size = 1024;
string d(size, 0);
generate(d.begin(), d.end(), [&] { return rd(); });
  • 类cmp函数的定义,适用 lower_bound()方法
    • 两个const都不能掉!
class typeUnassembled {
  public:
    size_t index;
    std::string data;
    typeUnassembled(size_t _index, std::string _data) : index(_index), data(_data) {}
    bool operator<(const typeUnassembled &t1) const { return index < t1.index; }
};
  • urg = static_cast<bool>(fl_b & 0b0010'0000); // binary literals and ' digit separator since C++14!!!

Computer Networking Lecture CS144 Stanford

Stanford CS144

[toc]

1-0 The Internet and IP Introduction

internet layer: Internet Protocol, IP address, packet's path

彩蛋:世一大惺惺相惜

Stanford-THU

用ping和traceroute看IP地址; 光纤2/3光速,8637km -> RTT=86ms

It's the Latency, Stupid
  • The distance from Stanford to Boston is 4320km.
  • The speed of light in vacuum is 300 x 10^6 m/s.
  • The speed of light in fibre is roughly 66% of the speed of light in vacuum.
  • The speed of light in fibre is 300 x 10^6 m/s * 0.66 = 200 x 10^6 m/s.
  • The one-way delay to Boston is 4320 km / 200 x 10^6 m/s = 21.6ms.
  • The round-trip time to Boston and back is 43.2ms.
  • The current ping time from Stanford to Boston over today's Internet is about 85ms:
  [cheshire@nitro]$ ping -c 1 lcs.mit.edu
  PING lcs.mit.edu (18.26.0.36): 56 data bytes
  64 bytes from 18.26.0.36: icmp_seq=0 ttl=238 time=84.5 ms
  • So: the hardware of the Internet can currently achieve within a factor of two of the speed of light.
1-1 A day in the life of an application
  • Networked Applications: connectivity, bidirectional and reliable data stream
  • Byte Stream Model: A - Internet - B, server和A、B均可中断连接
  • World Wide Web (HTTP: HyperText Transfer Protocol)
    • request: GET, PUT, DELETE, INFO, 400 (bad request)
    • GET - response(200, OK) , 200代表有效
    • document-centric: "GET/HTTP/1.1", "HTTP/1.1 200 OK \"
  • BitTorrent: peer-to-peer model
    • breaks files into "pieces" and the clients join and leave "swarms" of clients
    • 先下载 torrent file -- tracker 存储 lists of other clients
    • dynamically exchange data
  • Skype: proprietary system, a mixed system
    • two clients: A -- (Internet + Rendezvous server) -- NAT -- B
  • NAT(Network Address Translator): 连接的单向性,使得A只能通过Rendezvous server询问B是否直连A =>reverse connection
  • Rendezvous server
  • 如果模式是A -- NAT-- (Internet + Rendezvous server) -- NAT -- B,Skype用Relay来间接传递信息
Lego TCP/IP
1-2 The four layer Internet model
4-layer

4 layer: 利于reuse

Internet: end-hosts, links and routers

  • Link Layer: 利用 link 在 end host和router 或 router和router之间 传输数据, hop-by-hop逐跳转发
    • e.g. Ethernet and WiFi
  • Network Layer: datagrams, Packet: (Data, Header(from, to))
    • packets可能失去/损坏/复制,no guarantees
    • must use the IP
    • may be out of order
  • Transport Layer: TCP(Transmission Control Protocol) 负责上述Network层的局限性,controls congestion
    • sequence number -> 保序
    • ACK(acknowledgement of receipt),如果发信人没收到就resend
    • 比如视频传输不需要TCP,可以用UDP(User Datagram Protocol),不保证传输
  • Application Layer

two extra things

  • IP is the "thin waist" ,这一层的选择最少
  • the 7-layer OSI Model
7-layer
1-3 The IP Service
  • Link Frame (IP Datagram(IP Data(Data, Hdr), IP Hdr), Link Hdr )
  • The IP Service Model的特点
    • Datagram: (Data, IP SA, IP DA),每个 router 有 forwarding table,类比为 postal service 中的 letter
    • Unreliable: 失去/损坏/复制,保证只在必要的时候不可靠(比如queue congestion)
    • Best-effort attempt
    • Connectionless : no per-flow state, mis-sequenced
  • IP设计简单的原因
    • minimal, faster, streamlined
    • end-to-end (在end points implement features)
    • build a variety of reliable/unreliable services on top
    • works over any link layer
  • the IP Service Model
    1. tries to prevent packets looping forever (实现:在每个datagram的header加hop-count field: time to live TTL field, 比如从128开始decrement)
    2. will fragment packets if they're too long (e.g. Ethernet, 1500bytes)
    3. header checksum:增强可靠性
    4. allows for new versions of IP
    5. allows for new options to be added to header (由router处理新特性,慎重使用)
1-4 A Day in the Life of a Packet
  • 3-way handshake
    1. client: SYN
    2. server: SYN/ACK
    3. client: ACK
  • IP packets
    • IP address + TCP port (web server通常是80)
    • hops, Routers: wireless access point (WiFi的第一次hop)
    • forwarding table
    • default router
1-5 Principle: Packet switching principle

packet: self-contained

packet switching: independently for each arriving packet, pick its outgoing link. If the link is free, send it. Else hold the packet for later.

source packet: (Data, (dest, C, B, A)) 发展成只存destination,每个switch有table

two consequences

  • simple packet forwarding: No per-flow state required,state不需要store/add/remove
  • efficient sharing of links: busty data traffic; statistical multiplexing => 对packet一视同仁,可共享links
1-6 Principle: Layering
  • 一种设计理念,layers are functional components, they communicate sequentially
  • edit -> compile -> link -> execute
  • compiler: self-contained, e.g. lexical analysis, parsing the code, preprocessing declarations, code generation and optimization
  • layering的原因:1.modularity 2.well defined service 3.reuse 4.separation of concerns 5.continuous improvement 6.p2p communications
1-7 Principle: Encapsulation
  • TCP segment is the payload of the IP packet. IP packet encapsulates the TCP segment.
  • 一层层,套footer和header
* 两种写法,底层的写法(switch design)header在右边,software的写法(protocol)header在左边(IETF)
* VPN: (Eth, (IP, (TCP, (TLS, IP Packet)))),外层的TCP指向VPN gateway
1-8 Byte Order
  • 2^32 ~ 4GB ~ 0x0100000000
  • 1024=0x0400 大端:0x04 0x00;小端: 0x00 0x04.
  • Little endian: x86, big endian: ARM, network byte order
  • e.g. uint16_t http_port=80; if(packet->port==http_port){...} IPv4的packet_length注意大小端
  • 函数:htons(),ntohs(),htonl(),ntohl()
    • host/network, short/long
    • #include<arpa/inet.h>
1-9 IPv4 addresses

goal:

  • stitch many different networks together
  • need network-independent, unique address

IPv4:

  • layer 3 address
  • 4 octets a.b.c.d
  • 子网掩码netmask: 255.128.0.0 前9位,1越少网络越大,same network不需要路由,直接link即可
IPv4 Datagram

IPv4 Datagram

  • Total Packet Length: 大端,最多65535bytes, 1400 -> 0x0578
  • Protocol ID: 6->TCP

Address Structure

  • network+host
  • class A,B,C: 0,7+24; 10, 14+16; 110, 21+8

Classless Inter-Domain Routing (CIDR,无类别域间路由)

  • address block is a pair: address, count
  • counts是2的次方? 表示netmask长度
  • e.g. Stanford 5/16 blocks 5*2^(32-16)
  • 前缀聚合,防止路由表爆炸
  • IANA(Internet Assigned Numbers Authority): give /8s to RIRs
1-10 Longest Prefix Match(LPM)

forwarding table: CIDR entries

  • LPM的前提是必须先match,再看prefix
  • default: 0.0.0.0/0
1-11 Address Resolution Protocol(ARP)

IP address(host) -> link address(Ethernet card, 48bits)

Addressing Problem: 一个host对应多个IP地址,不容易对应

  • 解决方案:gateway两侧ip地址不同,link address确定card,network address确定host
  • 这有点历史遗留问题,ip和link address的机制没有完全地分离开,decoupled logically but coupled in practice
  • 对于A,ip的目标是B,link的目标是gateway

ARP,地址解析协议:由IP得到MAC地址 => 进一步可得到gateway address

  • 是一种request-reply protocol
  • nodes cache mappings, cache entries expire
  • 节点request a link layer broadcast address,然后收到回复,回复的packet有redundant data,看到它的节点都能生成mapping
  • reply:原则上unicast,只回传给发送者=>实际实现时更常见broadcast
  • No "sharing" of state: bad state will die eventually
  • MacOS中保留20min
  • gratuitous request: 要求不存在的mapping,推销自己
ARP

e.g.

hardware:1(Ethernet)

protocol: 0x0800(IP)

hardware length:6 (48 bit Ethernet)

protocol length:4(32 bit IP)

opcode: 1(request) /2(reply)

Destination: broadcast (ff:ff:ff:ff:ff:ff)

1-12 recap
1-13 SIP, Jon Peterson Interview

the intersection between technology and public policy

  • IETF ( The Internet Engineering Task Force)
  • ICANN(The Internet Corporation for Assigned Names and Numbers)

SIP(Session Initiation Protocol,会话初始协议)

  • end-to-end的设计
  • soft switching: 将呼叫控制功能从传输层分离
  • PSTN ( Public Switched Telephone Network ) -> VOIP(Voice over Internet Protocol): telephony replacement

SIP的应用场景

  • Skype内部协议转换成SIP
  • VOIP, FiOS( a telecom service offered over fiber-optic lines)

现代技术

  • SDN (Software Defined Network)
  • I2RS(interface to the routing system)
  • CDN(Content Delivery Network): 1.express coverage areas 2.advertise services that they provide, in order to allow collaboration or peering among CDNs => optimal selections of CDNs
  • 识别robo calling
2-0 Transport (intro)
  • 关注TCP的correctness
  • detect errors的三个算法:checksums, cyclic redundancy checks, message authentication codes
  • TCP(Transmission Control Protocol)、UDP(User Datagram Protocol)、ICMP(Internet Control Message Protocol)
2-1 The TCP Service Model

The TCP Service Model

  • reliable, end-to-end, bi-directional, in-sequence, bytestream service
    • Positive acknowledgement with retransmission
    • Peer TCP layers communicate: connection
    • 传输层方面,由于链路层带宽大增,TCP window scale option 被普遍使用,另外 TCP timestamps option 和 TCP selective ack option 也很常用
  • Flow control using sliding window (包括 Nagle 算法等)
    • 提高吞吐量,充分利用链路层带宽
    • tcp connection互不感知,缺少对网卡带宽的统筹安排
    • 原来设计 TCP 的时候,人们认为丢包通常是拥塞造成的,这时应该放慢发送速度,减轻拥塞;无线网中,丢包可能是信号太弱造成的,这时反而应该快速重试,以保证性能
  • congestion control
    • 防止过载造成丢包
    • 包括 slow start、congestion avoidance、fast retransmit 等

过程:三次握手和四次挥手(参考2-6的状态转移图理解)

Techniques to manufacture reliability

Remedies

  • Sequence numbers: detect missing data
  • Acknowledgments: correct delivery
    • Acknowledgment (from receiver to sender)
    • Timer and timeout (at sender)
    • Retransmission (by sender)
  • Checksums/MACs: detect corrupted data
    • Header checksum (IP)
    • Data checksum (UDP)
  • Window-based Flow-control: prevents overrunning receiver
  • Forward error correction (FEC)
  • Retransmission
  • Heartbeats

Correlated failure

TCP/DNS

Paradox of airplanes

The TCP Segment Format

TCP header
  • IANA port number: ssh 22, smtp 23, web 80
  • source port: 初始化用不同的port避免冲突
  • Flags
    • PSH flag: push,比如键盘敲击
    • URG应该在ACK前面
  • HLEN 和 (TCP options) 联系
TCP uniqueness 五个部分,104bit

唯一性

  • 要求source port initiator每次increment: 64k new connections
  • TCP picks ISN to avoid overlap with previous connection with same ID, 多一个域,增加随机性
  • ISN的意义在于:1)security,避免自己的window被overlap 2)便于filter out不同类型的包
2-2 UDP service model

不需要可靠性:app自己控制重传,比如早期版本的NFS (network file system)

UDP header * Checksum 对于 IPv4 可选,可以为全0 * Checksum 用了 IP header,违背 layering principle,是为了能detect错传 * UDP header 有 length 字段,而TCP没有,因为TCP对空间要求高,用隐含的方式计算 length * port demultiplexing, connectionless, unreliable

应用

DNS: domain name system,因为request全在单个datagram里

DHCP: Dynamic Host Configuration Protocol

  • new host在join网络时得到IP
  • 连WiFi

对重传、拥塞控制、in-sequence delivery 有 special needs 的应用,比如音频,但现在UDP不像以前用的那么多,因为很多是http,基于TCP。

2-3 The Internet Control Message Protocol (ICMP) Service Model

report errors and diagnoise problems about network layer

网络层work的三个因素:IP、Routing Tables、ICMP

ICMP

Message的意义见RFC 792

应用于ping:先发送8 0( echo request),再送回0 0(echo reply)

应用于traceroute:

  • 核心思想:连续发送TTL从1开始递增的UDP,期待回复的11 0(TTL expires)
    • Source is random and different for each; destination starts with a random number and increases by one for each
  • 由于路由选择问题,traceroute 无法保证每次到同一个主机经过的路由都是相同的。
  • traceroute 发送的 UDP 数据报端口号是大于 30000 的。如果目的主机没有任何程序使用该端口,主机会产生一个 3 3(端口不可达) ICMP报文给源主机。
2-4 End-to-End Principle

Why Doesn't the Network Help?

  • e.g.:压缩数据、Reformat/translate/improve requests、serve cached data、add security、migrate connections across the network
  • end-to-end principle: function的正确完整实现只依赖于通信系统的end points

end-to-end check

  • e.g. File Transfer: link layer的error detection只检测transmission错误,不检测error storage
  • e.g. TCP小概率会出错(stack)、BitTorrent
  • wireless link相比wire link功能复杂,可靠性低,所以在link layer重传,可提升TCP性能
  • RFC1958: "strong" end to end: 不推荐在 middle 实现任何功能,比如在 link layer 重传,假定了reliabilty的提升值得latency的牺牲
2-5 Error Detection: 3 schemes: 3 schemes
  • detect errors的三个算法:checksums, CRC(cyclic redundancy checks), MAC(message authentication codes)
  • 增补方式
    • append: ethernet CRC, TLS MAC
    • prepend: IP checksum
  • Checksum (IP, TCP)
    • not very robust, 只能检1位错
    • fast and cheap even in software
    • IP, UDP, TCP use one's complement算法:16-bit word packet求和,进位加到底部,再取反码(特例:0xffff -> 0xffff,因为在TCP,checksum field 为 0 意味着没有 checksum)
  • CRC: computes remainder of a polynomial (Ethernet),见通信与网络笔记
    • 通常是由网卡硬件完成的,在发包的时候由硬件填充 CRC,在收包的时候网卡自动丢弃 CRC 不合格的包
    • 虽然more expensive,但支持硬件计算
    • 可对抗2 bits error、奇数error、小于c bits的突发错(burst)
    • 可incrementally计算
    • e.g. USB(CRC-16): $\bf{M} = 0x8005 = x^{16}+x^{15}+x^2+1$,对于generator需要给左边pad 1
  • MAC: message authentication code: cryptographic transformation of data(TLS)
    • robust to malicious modifications, but not errors
    • 检错能力有局限,受随机性影响,不如CRC,no error detection guarantee
    • $c=MAC(M,s)$,M + c意味着对方有secret或者replay
    • 对于replay,ctr++, 具体见我的密码学笔记的TLS部分【目前尚未整理】
2-6 Finite State Machines
HTTP Request TCP Connection
  • 非常规路线的处理:比如对于第二个SYN或者FIN信号,接收机选择忽视,具体见bool TCPReceiver::segment_received(const TCPSegment &seg)的实现
2-7 Flow Control I: Stop-and-Wait
  • 核心是 receiver 给 sender 反馈,让sender不要送太多 packets
  • 基本方法
    • 方案一:stop and wait
    • 方案二:sliding window

stop and wait

  • flight 中最多一个 packet
  • 针对 ACK Delay(收到ACK的时间刚好在timeout之后)的情形,会有duplicates
    • 解决方案:用一个1-bit counter 提供信息
    • assumptions:1)网络不产生重复packets;2)不delay multiple timeouts
stop-and-wait
2-8 Flow Control II: Sliding Window
  • Stop-and-Wait的性能:RTT=50ms, Bottleneck=10Mbps, Ethernet packet length=12Kb => 性能(2%)远远不到瓶颈
  • Sliding Window计算Window size填满性能

Sliding Window Sender

  • Every segment has a sequence number (SeqNo)
  • Maintain 3 variables
    • Send window size(SWS)
    • Last acknowledgment(LAR)
    • Last segment sent(LSS)
  • Maintain invariant: $(LSS - LAR) \leq SWS$
  • Advance LAR on new acknowledgement
  • Buffer up to SWS segments

Sliding Window Receiver

  • Maintain 3 variables
    • Receive window size(RWS)
    • Last acceptable segment(LAS)
    • Last segment received(LSR)
  • Maintain invariant: $(LAS - LSR) \leq RWS$
  • 如果收到的packet比LAS小,则发送ack
    • 发送cumulative acks: 收到1, 2, 3, 5,发送3
    • TCP acks are next expected data,因此要加一,上个例子改为4,初值为0

RWS, SWS, and Sequence Space

  • $RWS \geq 1, SWS \geq 1, RWS \leq SWS$
  • if $RWS = 1$, "go back N" protocol ,need SWS+1 sequence numbers (需要多重传)
  • if $RWS = SWS$, need 2SWS sequence numbers
  • 通常需要$RWS+SWS$ sequence numbers:考虑临界情况,SWS最左侧的ACK没有成功发送,重传后收到了RWS最右侧的ACK

TCP Flow Control

  • Receiver advertises RWS using window field
  • Sender can only send data up to LAR+SWS
2-9 Retransmission Strategies

protocol可能的运转方式 (ARQ: automatic repeat request)

  • Go-back-N: pessimistic,重传ack, ack+1, ack+2 ...
    • e.g. RWS=1的情形
  • Selective repeat: optimistic,重传ack, last_sent, last_sent+1, ...
    • e.g. RWS=SWS=N的情形
    • 对burst of losses效果不好
2-10 TCP Header
TCP header
  • pseudo header:类似2-2,checksum的计算囊括了IP header
  • ack: 如果是bi-directional,也携带data信息;如果是uni-directional,好像不携带
  • URG: urgent, PSH: push
  • ACK: 除了第一个packet SYN,其它seg的ACK都置换为1
  • RST: reset the connection
  • urgent pointer:和URG联系,指出哪里urgent
2-11 TCP Setup and Teardown

状态机的实现很简洁,核心是如何 set up 和 clean up (port number, etc)

3-way handshake

Active opener and Passive opener

  1. client: SYN, 送base number(syqno) to identify bytes
  2. server: SYN+ACK, 也送base number
  3. client: ACK

支持“simultaneous open”

传送TCP segment,最小可以1byte,比如在ssh session打字

connection teardown

  1. client: FIN
  2. server: (Data +) ACK
  3. server: FIN
  4. client: ACK
Clean Teardown
  • 为什么 TCP 协议有 TIME_WAIT 状态
    • TIME_WAIT 仅在主动断开连接的一方出现,被动断开连接的一方会直接进入 CLOSED 状态,进入 TIME_WAIT 的客户端需要等待 2 MSL 才可以真正关闭连接
    • 不直接关闭连接的原因:
* 防止延迟的数据段被其他使用相同源地址、源端口、目的地址以及目的端口的 TCP 连接收到
  * RFC 793
  * `#define TCP_TIMEWAIT_LEN (60*HZ) /* how long to wait to destroy TIME-WAIT state, about 60 seconds	*/`
		* 但是如果主机在过去一分钟时间内与目标主机的特定端口创建的 TCP 连接数超过 28,232,那么再创建新的 TCP 连接就会发生错误,也就是说如果我们不调整主机的配置,那么每秒能够建立的最大 TCP 连接数为 ~470
* 保证 TCP 连接的远程被正确关闭,即等待被动关闭连接的一方收到 `FIN` 对应的 `ACK` 消息
  * 防止TIME-WAIT 较短导致的握手终止,服务端发送`RST`
  • 处理方案:除了上图的两者,还可以:
* 修改 `net.ipv4.ip_local_port_range` 选项中的可用端口范围,增加可同时存在的 TCP 连接数上限;
* Servers with high connection/transaction rates
  * TCP servers, e.g. web server
  * UDP servers, e.g. DNS server
* On multi-core systems, using multiple servicing threads, e.g. one thread per servicing core.
  * The single server socket becomes bottleneck
  * Cache line bounces
  * Hard to achieve load balance
  * Things will only get worse with more cores
  • Single TCP Server Socket
* solution 1: Use a listener thread to dispatch established connections to server threads
  * The single listener thread becomes bottleneck due to high connection rate
  * Cache misses of the socket structure
  * Load balance is not an issue here
* solution 2: All server threads accept() on the single server socket
  * Lock contention on the server socket
  * Cache line bouncing of the server socket
  * Loads (number of accepted connections per thread) are usually not balanced 
    * Larger latency on busier CPUs
    * It can almost be achieved by accept() at random intervals, but it is hard to decide the interval value, and may introduce latency
  • Single UDP Server Socket
  • New Socket Option - SO_REUSEPORT
* Allow multiple sockets bind()/listen() to the same local address and TCP/UDP port 
  * Every thread can have its own server socket
  * No locking contention on the server socket
* Every thread can have its own server socket No locking contention on the server socket
* Load balance is achieved by kernel - kernel randomly picks a socket to receive the TCP connection or UDP request
* For security reason, all these sockets must be opened by the same user, so other users can not "steal" packets
  • How to enable?
* sysctl net.core.allow_reuseport=1
* Before bind(), setsockopt SO_REUSEADDR and SO_REUSEPORT
* Then the same as a normal socket - bind()/listen() /accept()
  • Known Issues
* Hash
* Have not solved the cache line bouncing problem completely
  * Solved: The accepting thread is the processing thread
  * Unsolved: The processed packets can be from another CPU
    * Instead of distribute randomly, deliver to the thread/socket on the same CPU (input queue和server thread一一对应)
    * But hardware may not support as many RxQs as CPUs
* Some scheduler mechanism may harm the performance
  * Affine wakeup - too aggressive in certain conditions, causing cache misses
2-12 TCP Recap

IP和UDP都是best-effort and unreliable,但是我们不需要担心truncation和corruption,因为:

  • Header checksum (IP)
  • Data checksum (UDP)
HTTP HTTP3
2-13 TCP/IP -- Kevin Fall

《TCP/IP Illustrated》2nd edition

securites: firewalls; architectural underpinnings

packets和datagrams是两个核心概念,datagrams为了明确目的地,在设计时有更多的trade-off

3-d printing、枪、DRM(Digital Rights Management)

3-0 Packet Switching

Packet -> self-contained data unit

packet delay

  • Packetization delay
  • Propagation delay
  • Queueing delay
3-1 The History of Networks

Semaphore telegraphs by Chappe (France),发展出以下概念:

  • Codes
  • Flow Control
  • Synchronization
  • Error detection and retransmission
  • Encryption

Pre-defined messages -> arbitrary messages -> compression -> control signals "Protocols"

3-2 What is packet switching?

Circuit Switching

  • telephone: dedicated wire -> circuit switch -> dedicated wire
  • each phone call: 64 kb/s, no share with anybody else (private, guaranteed, isolated data rate from e2e)
  • A 10Gb/s trunk line can carry over 150000 calls

Circuit Switching 用于 Internet 的缺点

  • Inefficient: bursty communication (images, ssh connection, web pages)
  • Diverse Rates
  • State Management

Packet Switching

  • Network = end hosts + links + packet switches
  • forwarding table (routed individually by looking up)
    • All packets share the full capacity of a link
    • The routers maintain no per-communication state
  • have buffers: must send one at a time during periods of congestion
  • 有不同 types: routers、ethernet switches

Why Internet uses packet switching

  • Efficient use of expensive links
  • Resilience to failure of links & routers
    • the Internet was to be a datagram subnet
  • Internet was designed to be the interconnection of the existing networks
3-3 Terminology, End to End Delay and Queueing

Propagation Delay: $t_l = \frac{l}{c}$

  • single bit to travel over a link
  • 1000km, 2*10^8m/s ---> 5ms
  • 不受 link rate 影响

Packetization Delay: $t_p=\frac{p}{r}$

  • 64byte packet, 100Mb/s link ---> 5.12us
  • 1kbit (1024bit) packet, 1kb/s link (1000bit/s) ---> 1.024s

E2E delay $t=\sum_i(\frac{p}{r_i}+\frac{l_i}{c}+Q_i(t))$

  • store and forward network
  • router 理论上能等到 header 直接开始 packetization (cut through switching),internet router 通常不这样做,是收到整个 packet 再发送
  • queueing delay -> packet delay variation
3-4 Playback Buffers

Real-time applications (e.g. YouTube and Skype) have to cope with variable queueing delay

playback-buffer

  • variable delay 有下界
  • receive 曲线斜率有上界
3-5 Simple Deterministic Queue Model

$Q(t) = A(t)-D(t)$

d(t): 水平截距的差,表示单个 byte 的 queueing time

Q: Why not send the entire message in one packet?

A: parallel transmission across all links -> reduce e2e latency

---> Statistical Multiplexing Gain = 2C/R

3-6 Queueing Model Properties

Queues with Random Arrival Processes (Queueing Theory)

  • Bustiness increases delay
  • Determinism minimizes delay
  • Little's Result
    • $L=\lambda d$, where d = average delay, lambda = arival rate, L = average number that are in the queue
  • The M/M/1 queue
    • 用 Poisson process 建模 aggregation of many independent random events,lambda = arrival rate
    • network traffic is very bursty => 用 poisson 过程建模 the arrival of new flows
    • M/M/1 Queue: $d=\frac{1}{\mu-\lambda}, L=\lambda d = \frac{\frac{\lambda}{\mu}}{1-\frac{\lambda}{\mu}}$
3-7 Switching and Forwarding
Congestion Control
  • Why
    • What if the receiver’s window size is really big?
* Sender transmits too many segments. Most overflow router’s queue and are dropped. We call this “congestion.”
* Sender must resend the same bytes again and again. Eventually, stream comes out of receiver’s TCP correctly
  • The problem with unlimited sending: collapse and fairness
  • In networking, almost any problem that involves decentralized resource allocation = congestion control.
* receiver’s window (advertised from receiver to sender)
* “congestion window” cwnd (maintained by sender)
  • How much data can be “on the link” at any moment?
* (5 Mbit/s) x (100 ms) = 62.5 kilobytes
  • Ideal total number of bytes outstanding = bandwidth x delay product (BDP).
  • “No loss” window: anything less than BDP + max queue size.
  • Note: 用 window 不用 rate,误差小
* On each byte acknowledged: cwnd += (segment size)/cwnd
  • On loss, assume congestion. Cut cwnd in half!
* Loss inferred when:
  * segment was sent a long time ago, still not acknowledged
  * or several later-sent segments have been acknowledged
  • Slow-start: exponential growth at the beginning
* On each byte acknowledged: cwnd++
* On first loss, cut cwnd in half and revert to AIMD
  • rpc框架congestion control可能和tcp congestion control相结合
    • https://capnproto.org/news/2020-04-23-capnproto-0.8.html
    • it queries the send buffer size of the underlying network socket, and sets that as the “window size” for each stream.
    • But, the TCP socket buffer size only approximates the BDP of the first hop. A better solution would measure the end-to-end BDP using an algorithm like BBR.
  • Oracle STREAMS's Flow Control
    • 状态从后往前propagation的设计,canputnext()
    • 阻塞则 putbq
  • tcp协议栈默认关闭nodelay的
  • ```
if there is new data to send
  if the window size >= MSS and available data is >= MSS
    send complete MSS segment now
  else
    if there is unconfirmed data still in the pipe
      enqueue data in the buffer until an acknowledge is received
    else
      send data immediately
    end if
  end if
end if
  





### MLSys + Network

* [腾讯的ETH-X和阿里的Alink两者本质区别是什么?OISA的态度是什么?-三大scale-up网络标准齐聚ODCC 2024](https://zhuanlan.zhihu.com/p/720064056)



### potpourri

#### RFC

* RFC 792: ICMP Message
* RFC 821: SMTP
* [RFC 1958](https://datatracker.ietf.org/doc/rfc1958/?include_text=1):Architectural Principles of the Internet
* [RFC 2606](https://datatracker.ietf.org/doc/rfc2606/): localhost
* [RFC 6298](https://datatracker.ietf.org/doc/rfc6298/?include_text=1): Computing TCP's Retransmission Timer
* [RFC 6335](https://tools.ietf.org/html/rfc6335): port number
* RFC 7414: A Roadmap for TCP



* [TCP backlog: syns queue and accept queue](https://www.cnblogs.com/Orgliny/p/5780796.html)
* [What is a REST API?](https://www.youtube.com/watch?v=Q-BpqyOT3a8)
  * Representational State Transfer (REST)
  * Architecture style
  * Relies on a stateless, client-server protocol, almost alwasys HTTP
    * GET: retrieve data from a specified resource
    * POST: submit data to be processed to a specified resource
    * PUT: update a specified resource
    * DELETE
    * HEAD: same as get but does not return a body
    * OPTIONS: return the supported HTTP methods
    * PATCH: update partial resources
  * Treats server objects as resources that can be created or destroyed
  * GitHub REST API: https://docs.github.com/en/rest
  * 推荐 Postman 工具
* [AF_INET域与AF_UNIX域socket通信原理对比](https://blog.csdn.net/sandware/article/details/40923491)

#### HTTP `401 Unauthorized` 与 `403 Forbidden`

> 参考:[RFC 9110:401](https://www.rfc-editor.org/rfc/rfc9110.html#name-401-unauthorized)、[RFC 9110:403](https://www.rfc-editor.org/rfc/rfc9110.html#name-403-forbidden)、[RFC 6750:Bearer Token Error Codes](https://www.rfc-editor.org/rfc/rfc6750.html#section-3.1)。

先区分两个问题:**Authentication(认证)回答“你是谁”**,**Authorization(授权)回答“你能否做这件事”**。

| 状态码 | 协议语义 | 常见原因 | 下一步 |
| --- | --- | --- | --- |
| `401 Unauthorized` | 请求缺少目标资源认可的有效认证凭证;名字虽叫 Unauthorized,实际更接近 **Unauthenticated** | 没带 token、token 过期/无效、签名错误 | 根据 `WWW-Authenticate` challenge 登录、刷新或更换凭证;不要用同一凭证盲重试 |
| `403 Forbidden` | 服务端理解请求,但拒绝执行 | 身份有效但权限/scope 不足,也可能是 IP、租户、资源策略或 WAF 拒绝 | 改权限、身份、资源或策略;原样重试通常无效 |

标准的 `401` response 必须带至少一个 `WWW-Authenticate` challenge:

http HTTP/1.1 401 Unauthorized WWW-Authenticate: Bearer realm="example"


OAuth Bearer Token 中,`invalid_token` 通常对应 `401`,`insufficient_scope` 对应 `403`。但现实服务不总严格遵守:有些系统会用 `403` 表达临时风控或限流;客户端应结合 response body 中的稳定错误码和服务文档判断,标准限流状态应优先使用 `429 Too Many Requests`。

两个边界容易记错:

- `403` 不保证服务端已经认证出具体用户,它只保证服务端拒绝请求;匿名访问被策略禁止也可能返回 `403`。
- 为避免泄露资源是否存在,服务端可以用 `404 Not Found` 隐藏本应返回的 `403`,所以 `404` 也不总能证明资源不存在。

排障时,`401/403` 往往反而说明 DNS、TCP、TLS、路由和 HTTP 服务已经打通,问题已经进入认证/授权或应用策略层;它们与连接超时、DNS 失败、TLS handshake 失败不是同一层故障。

c++ #include #include #include #include #include #include #include

#define UNIX_SOCK_PATH_MAX_LEN (sizeof(((struct sockaddr_un*)0)->sun_path)) #define COMMAND_MAX_LEN 64 //file name format is project_pid.sock #define SOCK_PATH_FORMAT "/dev/shm/project_%llu.sock"

char control_sock_path[UNIX_SOCK_PATH_MAX_LEN] = {'\0',};

int main(int argc, char *argv[]) { int fd = socket(AF_UNIX, SOCK_STREAM, 0);

if(fd < 0){
	printf("socket error\n");
}
snprintf(control_sock_path,
		UNIX_SOCK_PATH_MAX_LEN,
		UNIX_SOCK_PATH_FORMAT,
		(unsigned long long)atoll(argv[1]));
printf("%s\n", control_sock_path);
struct sockaddr_un un;
memset(&un, 0, sizeof(un));
strncpy(un.sun_path, control_sock_path, sizeof(un.sun_path));
un.sun_family = AF_UNIX;
socklen_t len = offsetof(struct sockaddr_un, sun_path) + strlen(un.sun_path);
if (connect(fd, (struct sockaddr *)&un, len) < 0){
	close(fd);
	printf("connect error\n");
}
process(fd);

}


#### 实时订阅:连接只是载体,可靠性来自游标与重放

> 参考:[WHATWG Server-sent events](https://html.spec.whatwg.org/multipage/server-sent-events.html)、[RFC 6202: Bidirectional HTTP](https://www.rfc-editor.org/rfc/rfc6202.html)、[RFC 8895: ALTO Incremental Updates Using SSE](https://www.rfc-editor.org/rfc/rfc8895.html)、[NGINX `proxy_buffering`](https://nginx.org/en/docs/http/ngx_http_proxy_module.html#proxy_buffering)。

“实时订阅”不是某一种协议,而是一种状态同步关系:客户端先声明关注的 topic / resource,服务端在状态变化时持续推送 event。长连接只降低了事件到达延迟;一个可恢复的订阅还需要:

text subscription = filter + ordered event stream + cursor + reconnect + replay/resync


- **filter**:客户端能看到哪些资源和事件,必须在服务端重新做鉴权,不能只信客户端传来的 topic。
- **event stream**:事件要有稳定 schema 和顺序语义;跨 partition 是否全序,必须明确。
- **cursor**:记录客户端已经处理到哪里,例如 SSE 的 `id` / `Last-Event-ID`。
- **reconnect**:连接断开后重新建立,并使用退避和 jitter 防止大规模同时重连。
- **replay / resync**:游标仍在保留窗口内就补发缺失事件;游标过旧或出现 gap,就重新拉 snapshot。

因此,SSE 的自动重连不等于可靠投递。若服务端只向当前连接写数据、没有 durable event log,断线期间的事件仍会丢失;若事件可能重放,客户端还要按 `event_id` 幂等应用。工程上通常追求 **at-least-once + idempotency**,不要轻易声称 exactly-once。

##### SSE 的协议语义

SSE(Server-Sent Events)是在一个长时间不结束的 HTTP response 中,由服务端持续发送 UTF-8 文本事件。这里的“长连接”更准确地说是长生命周期的 HTTP stream:HTTP/1.1 下通常占用一条连接,HTTP/2 / HTTP/3 下则可与其他 stream 复用底层连接。HTTP chunk / frame 只是传输分块,可能被中间层重组;SSE 的业务事件边界始终是空行。

浏览器原生客户端是 `EventSource`,数据格式为 `text/event-stream`:

text HTTP/1.1 200 OK Content-Type: text/event-stream Cache-Control: no-cache X-Accel-Buffering: no

event: progress id: 42 retry: 3000 data: {"task_id":"t1","percent":80}

: heartbeat


| 字段 | 语义 |
| --- | --- |
| `data` | 事件载荷;连续多个 `data:` 行会用换行符连接。 |
| `event` | 事件类型;缺省时触发 `message`。 |
| `id` | 更新客户端保存的 last event ID;重连时浏览器通过 `Last-Event-ID` 发回。 |
| `retry` | 建议的重连等待时间,单位为毫秒。 |
| `:` | 注释行,不触发业务事件,常用作应用层 heartbeat。 |

javascript const source = new EventSource("/api/tasks/t1/events");

source.addEventListener("progress", event => { const update = JSON.parse(event.data); renderProgress(update); });

source.onerror = () => { // EventSource 默认会重连;这里只做状态展示和观测,不要再开第二条连接。 };

// 页面或任务不再需要订阅时必须主动释放。 source.close();


原生 `EventSource` 的请求方向是 client -> server,业务数据方向只有 server -> client;构造器只暴露 URL 和 `withCredentials`,不能方便地携带 POST body 或自定义 `Authorization` header。浏览器场景通常使用同源 cookie、短期签名 URL,或改用基于 `fetch()` 的流式客户端。跨域订阅还要正确配置 CORS 与 credentials;不要把长期 token 放进 URL,因为 URL 容易进入日志和监控。

服务端可用 HTTP `204 No Content` 告诉原生客户端停止重连。普通断开会触发自动重连;生产实现还应发送周期性注释 heartbeat,避免代理、网关或负载均衡器把空闲连接回收。15 秒只是 WHATWG / RFC 示例中的经验值,实际间隔必须小于整条链路上最短的 idle timeout。

##### Snapshot + delta:避免订阅启动时的竞态

典型同步流程不是“先查一次、再随便开条 SSE”,而是:

text GET snapshot -> 返回 state + snapshot_cursor SUBSCRIBE after=snapshot_cursor -> replay(cursor, current] -> 持续接收 live events 发现 cursor 过期 / 序号跳跃 -> 丢弃局部推断,重新获取 snapshot


`snapshot_cursor` 把快照与增量流接起来,避免“读取快照之后、建立订阅之前”发生的更新落入缝隙。更严格的实现应保证 snapshot 对应一个确定的日志位置;否则即使有 cursor,也可能重复或漏掉边界事件。

##### 选型

| 机制 | 数据方向与状态 | 优点 | 适用场景 / 主要代价 |
| --- | --- | --- | --- |
| 短轮询 | client 定时 pull;每次独立 request | 最简单、易缓存、易降级 | 低频状态;延迟与空请求开销互相制约。 |
| 长轮询 | server 暂挂 request,有事件或超时才返回;客户端立即再请求 | 兼容普通 HTTP,天然以 response 分帧 | 低频通知、旧基础设施;每轮仍有完整 header 和重建请求的间隙。 |
| SSE | 单向 server -> client;一条流式 HTTP response | 浏览器原生、文本分帧、自动重连、支持 event ID | 通知、任务进度、日志、feed;不适合高频双向交互和二进制流。 |
| WebSocket | 全双工长连接;应用自定义 message protocol | 双向、低开销、支持二进制 | 聊天、协同编辑、控制面;重连、恢复、鉴权续期和心跳都要自行设计。 |
| Webhook | server -> server 的独立 HTTP callback | 不要求订阅方维持连接,适合系统集成 | 延迟通常较高;要做签名、重试、去重和死信处理。 |

LLM API 常说“用 SSE 流式返回 token”,很多实现实际是 **SSE 格式的 POST streaming response**,客户端用 `fetch()` / SDK 逐块解析;它不一定使用浏览器原生 `EventSource`,`[DONE]` 等结束标记也属于应用协议,不是 SSE 标准字段。

##### 生产检查清单

- **代理缓冲**:应用必须及时 flush;NGINX 默认 `proxy_buffering on`,可对该路由关闭,或由 upstream 返回 `X-Accel-Buffering: no`。否则服务端明明逐条写,客户端却成批收到。
- **超时与心跳**:核对 browser、CDN、WAF、LB、gateway、reverse proxy、server 各层 idle / read timeout;heartbeat 要穿过整条链路。
- **慢消费者与背压**:为每个 subscriber 使用有界队列,明确 `drop / coalesce / disconnect / resync` 策略。TCP 变慢只会把压力向上游传播,不会替应用决定保留哪些业务事件。
- **容量**:长连接主要消耗 file descriptor、socket / request state、内存和负载均衡连接槽;关注 `active_connections`、连接建立率、重连率、发送队列大小、event lag、drop / replay 数和连接时长。
- **生命周期**:客户端切换资源、页面卸载或任务结束时主动 `close()`;服务端检测断连并取消 producer,避免后台继续计算和写入。
- **正确性测试**:覆盖断网重连、重复事件、乱序 / gap、游标过期、代理缓冲、token 过期、服务重启和重连风暴,而不只测正常持续输出。

#### Proxy / Tunnel / SSH Port Forwarding

参考:[代理,网关,隧道,有什么区别与联系? - 知乎](https://www.zhihu.com/question/268204483/answer/334644846)

##### Squid:可缓存、可治理的 HTTP 正向代理

[Squid](https://github.com/squid-cache/squid) 是应用层 Web proxy/cache。典型部署把它放在客户端出口:先执行 ACL、认证和路由,再直接访问 origin 或转发给 parent proxy;对可缓存 HTTP 响应,还会按新鲜度和再验证规则复用对象。

text client -> Squid

       |-- HIT  -> cached response
       `-- MISS -> origin / parent proxy -> cache if allowed -> client

它的能力可以拆成四组:

- **出口治理**:按来源、目标域名、端口和请求类型做 ACL,统一认证、访问日志与审计。
- **HTTP 缓存**:降低重复请求的延迟和出口带宽;缓存与代理彼此独立,也可以配置成只代理、不缓存。
- **代理层级**:通过 `cache_peer` 组织 parent / sibling cache,集中管理上游出口。
- **其他模式**:也能作为 reverse proxy 或 interception proxy,但不是理解 Squid 的首要入口。

HTTPS 默认通过 `CONNECT host:443` 建立 TCP tunnel。Squid 可以控制目标主机和端口,但看不到 TLS 内的 path、query、header 和 body,也就不能缓存或按内容治理。`SSL-Bump` 通过部署受信 CA 做 TLS 中间人才能重新获得这些能力,同时引入隐私、合规和证书安全风险,不应视为普通缓存配置。[HTTPS / CONNECT 边界](https://wiki.squid-cache.org/Features/HTTPS)

今天通用 Web 流量大量采用 HTTPS、动态内容和 CDN,Squid 的普适缓存收益弱于早期互联网。它仍适合需要**统一出口、访问策略、审计、parent routing**,或明确存在可缓存对象的受控网络。显式配置客户端使用 proxy,语义通常也比透明拦截更清楚;interception 会破坏端到端假设,并影响认证、协议兼容和故障定位。[Interception 的限制](https://wiki.squid-cache.org/SquidFaq/InterceptionProxy)

一次实战经验:远端 headless runtime 认证已经成功,但 `exec` 仍失败。最后发现问题不在 login,而在网络出口:远端能读本地 auth cache,却无法稳定访问运行时依赖的外部 endpoint。解决方式是本机启动 loopback HTTP CONNECT proxy,再用 SSH reverse tunnel 把远端 loopback 端口接到本机 proxy。

这类问题要先分清三层:

- **Proxy**:代理应用层请求。HTTP proxy 会理解 HTTP 请求;HTTPS 走 HTTP proxy 时通常用 `CONNECT host:443`,让 proxy 建立一条 TCP 隧道,之后 TLS 流量在隧道里透传。
- **Tunnel**:改变网络可达性,本质是把一个连接封装进另一条连接里。隧道不一定理解上层协议,只负责转发字节流。
- **Port forwarding**:端口级隧道。把一端的 `host:port` 映射到另一端的 `host:port`,SSH 只是最常见的安全承载层。

常见 SSH 转发模式:

bash # local forward:本机监听 18081,访问 remote 视角可达的 target:443。 ssh -N -L 127.0.0.1:18081:target.example.com:443 user@remote

# reverse forward:remote 监听 18081,回连本机 18080。 ssh -N -R 127.0.0.1:18081:127.0.0.1:18080 user@remote

# dynamic forward:本机启动 SOCKS 代理,目标地址由客户端请求动态决定。 ssh -N -D 127.0.0.1:1080 user@remote


其中 `-R` 最容易想反。它是在 **remote 机器上开 listener**,但每次 remote 有连接进来,SSH 会把连接沿着已经建立的 SSH 会话带回本机,再连到本机侧的目标地址。

本次拓扑可以抽象成:

text remote app -> HTTP_PROXY=http://127.0.0.1:18081 -> remote loopback listener -> SSH reverse tunnel -> local 127.0.0.1:18080 CONNECT proxy -> public internet / target endpoint


`127.0.0.1` 是关键安全边界。无论本机 proxy 还是 remote listener,默认都应绑定 loopback,而不是 `0.0.0.0`。前者避免把本机代理暴露给局域网或公网;后者避免把远端 tunnel 端口变成公开代理。

实战命令骨架:

bash # 1. 本机启动一个只监听 loopback 的 HTTP CONNECT proxy。 # 具体工具可替换,原则是 local 127.0.0.1:18080 提供 CONNECT 能力。

# 2. 建立 reverse tunnel:remote 18081 -> local 18080。 ssh -N \ -o ExitOnForwardFailure=yes \ -o ServerAliveInterval=30 \ -o ServerAliveCountMax=3 \ -R 127.0.0.1:18081:127.0.0.1:18080 \ user@remote

# 3. remote runtime 注入 proxy env。 export HTTP_PROXY=http://127.0.0.1:18081 export HTTPS_PROXY=http://127.0.0.1:18081 export ALL_PROXY=http://127.0.0.1:18081 export NO_PROXY=localhost,127.0.0.1


端口最好区分 remote listener 和 local upstream,例如 `remote 18081 -> local 18080`,不要偷懒写成同号端口。不同端口让排障语义清楚:`18081` 是远端入口,`18080` 是本机代理;也能减少端口复用、自引用、旧进程残留带来的误判。

验收要分层,不要只看“SSH 进程还在”:

bash # remote:listener 是否真的存在。 ss -ltnp | grep 18081

# local:proxy 是否真的在监听。 lsof -nP -iTCP:18080 -sTCP:LISTEN

# remote:通过 proxy 访问外部 endpoint。 HTTPS_PROXY=http://127.0.0.1:18081 \ curl -I --max-time 15 https://api.openai.com/v1/models

# 对照:不走 proxy 直连,判断是不是远端出口本身有问题。 curl -I --max-time 15 https://api.openai.com/v1/models


排障时要注意:`401 Unauthorized` 可能是健康信号。对需要鉴权的 API endpoint 来说,`401` 说明 TCP、TLS、DNS、路由都已经打通,只是业务凭证没带;超时、DNS 失败、TLS handshake 卡住才更像网络面问题。

这类远端 runtime 问题的通用 checklist:

- 先拆 surface:auth、network、entrypoint、runtime,不要把所有失败都归因到 login。
- 先探活 direct,再探活 proxy,比较错误形态。
- 明确目标域名;很多工具不只访问 `api.*`,还会访问 Web app/backend API、WebSocket、MCP endpoint。
- 确认工具是否真的读取 `HTTP_PROXY` / `HTTPS_PROXY` / `ALL_PROXY`;有些程序需要显式配置。
- wrapper 可以注入 proxy env,但不要在 wrapper 里保存 token。
- tunnel 只解决可达性,不解决账号权限、workspace policy、TLS 信任和业务鉴权。
- 用 `ExitOnForwardFailure=yes` 防止 SSH 看似成功但端口没开;长连再配合 keepalive 或 supervisor。

##### `ssh -R` 与 Unix socket reverse forward 的排障模型

参考:[ssh(1)](https://man.openbsd.org/ssh.1)、[ssh_config(5)](https://man.openbsd.org/ssh_config)、[sshd_config(5)](https://man.openbsd.org/sshd_config)、[unix(4)](https://man.openbsd.org/unix.4)。

一类常见故障:TCP reverse forward 单独可用,但 supervisor 同时拉起 TCP forward 和 Unix socket forward 时,`ssh -R` 直接以 `255` 退出,外层只看到 `tunnel exited`。这通常不是认证问题,而是某个 listener 没有 bind 成。

`ssh -R` 的语义是 remote 侧开 listener,本机侧接 upstream:

bash # TCP reverse forward: remote 127.0.0.1:18081 -> local 127.0.0.1:18080 ssh -N -R 127.0.0.1:18081:127.0.0.1:18080 user@remote

# Unix socket reverse forward: remote socket path -> local socket path ssh -N -R /tmp/remote-bridge.sock:/tmp/local-bridge.sock user@remote


TCP port 和 Unix-domain socket 的生命周期不一样:

- TCP listener 是内核里的 `ip:port` 绑定。进程退出后 listener 通常随之消失;残留问题更多表现为旧进程仍在监听、端口被占用、TIME_WAIT/复用策略干扰。
- Unix-domain socket 的地址是文件系统路径。`bind()` 会在文件系统里创建 socket 文件;socket 关闭后这个路径不会自动删除,必须显式 `unlink`。所以一次失败 run 留下的 `/tmp/*.sock`,就可能让下一次 `ssh -R remote_socket:local_socket` 直接 bind 失败。
- OpenSSH 有 `StreamLocalBindUnlink=yes`,用于创建 Unix-domain socket 前删除已有 socket 文件。但它不是可以无脑依赖的全局垃圾回收:客户端和服务端都有相关配置入口,实际是否生效取决于谁在创建这个 socket、命令是否经过跳板、多跳链路是否把 option 传到正确一端、远端权限是否允许删除。

排查顺序要把层拆开:

text

  1. 先测 TCP reverse forward 如果 TCP 都不通,优先查 SSH 参数、跳板、GatewayPorts、remote bind address、端口占用。
  1. 再测 Unix socket reverse forward 如果 TCP 通而 socket 不通,优先查 remote socket path 是否残留、目录权限、StreamLocalBindUnlink 是否命中实际 bind 方。
  1. 最后复现 supervisor 的完整 tunnel command 如果单测都通而完整命令失败,再查 supervisor 是否并发创建多个 forward、是否复用旧 remote path、是否正确清理失败 run。

更好的工程修法:不要把 `rm -f /tmp/xxx.sock` 留给启动脚本或人工操作,而要让拥有 tunnel 生命周期的 supervisor 负责。

text tunnel supervisor invariant: preflight:

- 远端 socket path 由 supervisor 生成和持有
- 启动 ssh -R remote_socket:local_socket 前,先通过 SSH 清理自己持有的旧 remote socket

start:

- 使用 ExitOnForwardFailure=yes,让 forward 没建成时立即失败
- 区分 TCP forward smoke 和 Unix socket forward smoke

observe:

- 私有日志记录真实 remote path 和 ssh stderr
- public payload 只暴露 cleanup succeeded/failed、tunnel exit code、smoke result

背后的通用知识点是:**长程 agent / benchmark runner 里的 tunnel、socket、lock file、pid file 都是有生命周期的资源,不是一次性命令字符串。** 谁拥有资源,谁就要负责 preflight cleanup、idempotent start、smoke test、失败证据和脱敏输出。否则一次失败 run 留下的状态会污染下一次 run,表现成“明明什么都没改,重跑又坏了”。

#### wireshark
[谈谈Linux中的TCP重传抓包分析](https://segmentfault.com/a/1190000019734707)

telnet cs144.keithw.org http GET /hello HTTP/1.1 # path part,第三个slash后面的部分 Host: cs144.keithw.org # host part,https://和第三个slash之间的部分

tcp.port == 90 and ip.addr== XXX tcp.len > 0 ip.ttl == XXX icmp.code == 0


课程作业:

1.Ping

2.SMTP:在TCP上层

3.Traceroute
* VM的第一跳是到laptop,不会decrement the TTL,因此hop 10对应TTL 9

#### curl
`curl` 是一个强大的命令行工具,用于通过URL进行数据传输。其原理可以看作是应用层和传输层协议的完整命令行实现:
* **DNS解析**: 将URL中的主机名(如 `cs144.keithw.org`)解析为IP地址。
* **建立TCP连接**: 与目标服务器的指定端口(HTTP为80,HTTPS为443)进行TCP三次握手,建立连接。
* **(HTTPS)TLS握手**: 如果是HTTPS请求,会在TCP连接之上进行TLS握手,协商加密算法,建立安全的加密通道。
* **发送HTTP请求**: 构造并发送一个HTTP请求报文。最简单的 `curl http://example.com` 会发送一个 `GET / HTTP/1.1` 请求,并附带 `Host: example.com` 等头部信息。
* **接收HTTP响应**: 读取服务器返回的HTTP响应报文,包括状态码(如 `200 OK`)、响应头和响应体(即HTML页面内容、JSON数据等)。
* **输出**: 默认情况下,`curl` 会将响应体打印到标准输出。
* **关闭连接**: 完成数据传输后,关闭TCP连接。

`curl` 支持众多协议(HTTP, HTTPS, FTP, SCP等)和复杂操作(如POST数据、设置Header、处理Cookie),是网络调试和自动化脚本中不可或缺的工具。


#### WiFi 与 路由器

* WIFI5 的连接速度最高 866.7 Mbps,只有开启 WIFI6 模式,并且启用160MHZ,才能突破 866.7 Mbps
* 路由器 LAN-LAN 级联
* 路由器的延时问题
  * One major modem manufacturer has contacted me, and we've been investigating where the time goes. It seems that there is room for improvement, but unfortunately modems will never be able to match ISDN. The problem is that over a telephone line, electrical signals get "blurred" out. In order to decode just one single bit, a 33.6kb/s modem needs to take not just a single reading of the voltage on the phone line at that instant, but that single reading plus another 79 like it, spaced 1/6000 of a second apart. A mathematical function of those 80 readings gives the actual result. This process is called **"line equalization"**. Better line equalization allows higher data rates, but the more "taps" the equalizer has the more delay it adds. The V.34 standard also specifies particular scrambling and descrambling of the data, which also take time. According to this company, the theoretical best round-trip delay for a 14.4kb/s modem (no compression or error recovery) should be 40ms, and for a 33.6kb/s modem 64ms. The irony here is that as the capacity goes up, the best-case latency gets worse instead of better. **For a small packet, it would be faster for your modem to send it at 9.6kb/s than at 33.6kb/s!**


#### TCP 吞吐研究

* 千兆以太网的裸吞吐量是 125MB/s,应用层的吞吐率大约在 117MB/s 上下
  * 【2022年】普遍机器带宽是25or100GB/s了,量级和内存带宽接近,极限情况下的包相关memcpy不可忽略
  * 在不考虑 jumbo frame 的情况下,计算过程是: 对于千兆以太网,每秒能传输 1000Mbit 数据,即 125000000B/s,每个以太网 frame 的固定开销有:preamble(8B)、MAC(12B)、type(2B)、payload (46B ~ 1500B)、CRC(4B)、gap(12B),因此最小的以太网帧是 84B,每秒可发送约 1488000 帧(换言之,对于一问一答的 RPC、其 qps 上限约是 700k/s),最大的以太网帧是 1538B,每秒可发送 81 274 帧。 
  * 再来算 TCP 有效载荷:一个 TCP segment 包含 IP header(20B)和 TCP header(20B),还有 Timestamp option(12B),因此 TCP 的最大吞吐量是 81274 × (1500-52) = 117MB/s,合 112MiB/s。


#### 常见高危端口

在网络安全中,某些端口因其关联的服务非常核心或存在固有弱点,而成为攻击者的重点扫描和攻击目标。

*   **21 (FTP - 文件传输协议)**: 主要风险在于其默认以**明文传输**数据和用户凭证,极易被网络嗅探工具截获。
*   **22 (SSH - 安全外壳协议)**: 协议本身安全,但作为服务器远程管理的主要入口,是**暴力破解攻击**的常见目标。安全策略包括:禁用密码登录(改用密钥)、禁止root直接登录、更改默认端口。
*   **3389 (RDP - 远程桌面协议)**: Windows远程桌面服务的默认端口。因其广泛使用和通常具备高权限,是勒索软件和黑客攻击的重点目标。
*   **3306 (MySQL)**: MySQL数据库服务的默认端口。直接暴露在公网是极大的安全隐患,容易导致数据泄露或被攻击。#

### 论文阅读

#### 《Ethane: Taking Control of the Enterprise, SIGCOMM 07》

make networks more manageable and more secure,一种思路是全方位的增加控制,相当于新增一层,只是hide了复杂度;于是提出ethane:

Ethane的思想

* The network should be governed by policies declared over high-level names
* Policy should determine the path that packets follow
* The network should enforce a strong binding between a packet
  and its origin.

Ethane的优势

* Security follows management.
* Incremental deployability.
* Significant deployment experience.

设计思想

* Controllers: 决定是否允许packet传输
* Switches: a simple flow table and a secure channel to the Controller
  * flow是一个重要的属性概念
  * Binding: When machines use DHCP to request an IP address, Ethane assigns it knowing to which switch port the machine is connected, enabling Ethane to attribute an arriving packet to a physical port.



其它细节

* Replicating the Controller: Fault-Tolerance and Scalability
  * cold-standby (having no network binding state) or warm-standby (having network binding state) modes

### Computer-Networking-A-Top-Down-Approach

#### 0.资料

* [Stanford CS144](https://github.com/huangrt01/Markdown-Transformer-and-Uploader/blob/master/Notes/Output/Stanford-CS144.md)

* [质量最高的pdf](https://download.csdn.net/download/mnt248/10350433)

* 习题
  * https://github.com/moranzcw/Computer-Networking-A-Top-Down-Approach-NOTES
  * https://github.com/HanochShi/Supplements-ComputerNetworking-ATopDownApproach-7th-ed
  * https://github.com/myk502/Top-Down-Approach


---

## Algorithms Potpourri

## Algorithms-Potpourri

* [格林公式在面积并问题中的应用](https://trinkle23897.github.io/posts/calc-circle-area-union) --- by n+e

* [Three Optimization Tips for C](https://www.slideshare.net/andreialexandrescu1/three-optimization-tips-for-c-15708507)

  * You can't improve what you can't measure

  * Reduce strength

  * Minimize array writes

c++ uint32_t digits10(uint64_t v) { if (v < P01) return 1; if (v < P02) return 2; if (v < P03) return 3; if (v < P12) {

if (v < P08) {
  if (v < P06) {
    if (v < P04) return 4;
    return 5 + (v < P05);
  }
  return 7 + (v >= P07);
}
if (v < P10) {
  return 9 + (v >= P09);
}
return 11 + (v >= P11);

} return 12 + digits10(v/P12); }

unsigned u64ToAsciiTable(uint64_t value, char* dst) { static const char digits[201] =

"0001020304050607080910111213141516171819"
"2021222324252627282930313233343536373839"
"4041424344454647484950515253545556575859"
"6061626364656667686970717273747576777879"
"8081828384858687888990919293949596979899"

uint32_t const length = digits10(value); uint32_t next = length - 1; while (value >= 100) {

auto const i = (value % 100) * 2;
value /= 100;
dst[next] = digits[i + 1];
dst[next - 1] = digits[i];
next -= 2;

} if (value < 10) {

dst[next] = '0' + uint32_t(value);

} else {

auto i = uint32_t(value) * 2;
dst[next] = digits[i + 1];
dst[next - 1] = digits[i];

} return length; }


### 数据

在GPU里针对单精度和双精度就需要各自独立的计算单元, 一般在GPU里支持单精度运算的单精度ALU(算术逻辑单元)称之为FP32 core, 而把用作双精度运算的双精度ALU称之为DP unit或者FP64 core. 

- FP64: 用8个字节来表达一个数字, 1位符号, 11位指数, 52位小数,**有效位数为16位**. 常用于科学计算, 例如: 计算化学, 分子建模, 流体动力学
- FP32: 用4个字节来表达一个数字, 1位符号, 8位指数, 23位小数,**有效位数为7位**. 常用于多媒体和图形处理计算、深度学习、人工智能等领域
- FP16: 用2个字节来表达一个数字, 1位符号, 5位指数, 10位小数,**有效位数为3位**. 常用于精度更低的机器学习等



### 数据结构

#### Hash Table

* Open Hashing (Closed Addressing) v.s. Close Hashing (Open Addressing)
  * https://programming.guide/hash-tables-open-vs-closed-addressing.html
  * Open Hashing 的缺点:
    * 读放大
    * load key 有额外一次 memory read
    * rehash 时整个结构要重建
* [Hopscotch hashing](https://en.wikipedia.org/wiki/Hopscotch_hashing)
* Cuckoo Hashing: https://web.stanford.edu/class/archive/cs/cs166/cs166.1146/lectures/13/Small13.pdf
* GeoHash
* Key的均衡性
  * [MurmurHash](https://github.com/aappleby/smhasher/blob/master/src/MurmurHash3.cpp)
* HashTable实现中可用的设计:
  * per dim合并
    * 记录available indices提升分配速度
  
  * map只存indices,提升rehash速度
  

##### Hash Collision

https://preshing.com/20110504/hash-collision-probabilities/

* 存在冲突的概率:$$\frac{k^2}{2N}$$

![image-20251028140434370](https://raw.githubusercontent.com/huangruiteng/CS-Notes/HEAD/Algorithms-Potpourri/image-20251028140434370.png)

* 冲突次数的期望:近似为 $$\frac{k^2}{N}$$
  * 冲突率:近似为 $$\frac{k}{N}$$
  * 证明:1)利用二项分布;2)假设每个位置最多冲突一次




#### LRU cache

* [linked-list based LRU cache](https://krishankantsinghal.medium.com/my-first-blog-on-medium-583159139237)
  * hashmap的value存双向链表节点的指针
* array-list based LRU cache
  * hashmap的value存array-list的index
    * Array-list的value存pre-index + post-index + entry
  * e.g. Persia

#### Radix Tree

[radix tree (Linux 内核实现)](https://lwn.net/Articles/175432/):压缩前缀树,维护 kv 查找

前缀树的结构:

* 一棵子树的所有子节点都有相同前缀的 key 值

* 只有叶子节点有对应的 value 值

* key的最大长度固定

当 key 有以下特性的时候压缩前缀树比 hashmap 更具优势:

* 当存储的 key 值本身就有很好的 hash 特性,但是又非常稀疏时。比如说网段,地址空间,可以不用设计复杂 hash 函数。每次插入查找没有hash 计算的开销。
* 大量的 key 有着相同的前缀时,相比于 hashmap 每个节点都要存储完整的 key 值,更具有空间复杂度优势。

应用:[Trie](https://en.wikipedia.org/wiki/Trie)

#### Pairing Heap

https://en.wikipedia.org/wiki/Pairing_heap

实现简单、均摊复杂度优越,用于实现优先队列

定义:一个配对堆要么是一个空堆,要么由一个根元素与一个可能为空的配对堆子树列表所组成。所有子树的根元素都大于该堆的根元素。

python type PairingTree[Elem] = Heap(elem: Elem, subheaps: List[PairingTree[Elem]]) type PairingHeap[Elem] = Empty | PairingTree[Elem]


操作:

* 合并:一个空堆与另一个堆合并将会返回另一个堆;否则将会返回一个新堆,其将两个堆的根元素中较小的元素当作新堆的根元素,并将较大的元素所在的堆合并到新堆的子堆中。

C++ function merge(heap1, heap2: PairingHeap[Elem]) -> PairingHeap[Elem] if heap1 is Empty

return heap2

elsif heap2 is Empty

return heap1

elsif heap1.elem < heap2.elem

return Heap(heap1.elem, heap2 :: heap1.subheaps)

else

return Heap(heap2.elem, heap1 :: heap2.subheaps)

* 插入:将一个仅有该元素的新堆与需要被插入的堆合并。

c++ function insert(elem: Elem, heap: PairingHeap[Elem]) -> PairingHeap[Elem] return merge(Heap(elem, []), heap)


* 删除最小:根元素即为最小元素。删除根元素,然后合并子树,合并方法为从左到右两两合并,然后再从左向右顺序合并。

c++ function delete-min(heap: PairingHeap[Elem]) -> PairingHeap[Elem] if heap is Empty

error

else

merge-pairs(heap.subheaps)
return elem: Elem

function merge-pairs(list: List[PairingTree[Elem]]) -> PairingHeap[Elem] if length(list) == 0

return Empty

elsif length(list) == 1

return list[0]

else

return merge(merge(list[0], list[1]), merge-pairs(list[2..]))

#### BloomFilter

https://en.wikipedia.org/wiki/Bloom_filter



### 数据处理

#### TopK

* std是对topK的近似

python def get_topk_amax(tensor, percentile):

  tensor = tf.abs(tensor)
  tensor = tf.reshape(tensor, [-1])
  tensor_size = tf.cast(tf.size(tensor), tf.float32)
  k = tf.math.maximum(1, tf.cast(tf.math.ceil(tensor_size * (1-percentile)), tf.int32))
  topk, _ = tf.math.top_k(tensor, k=k)
  target_amax = topk[-1]
  return target_amax

def get_std_amax(tensor, std_scale):

  return tf.abs(tf.reduce_mean(tensor)) + std_scale * tf.math.reduce_std(tensor)      

def get_max_amax(tensor):

  return tf.reduce_max(tf.abs(tensor))

```

图算法

本文由 GitVP 从 GitHub 收录并在站内全文呈现,版权归原作者所有(MIT)。

← 回到全部文章

同分类还有