COMSOL-并行計(jì)算

參考blog：https://www.comsol.com/blogs/added-value-task-parallelism-batch-sweeps/

我們知道并行計(jì)算可以加快計(jì)算速度，但是這個(gè)加快不是無限制的场航，而且這個(gè)速度的加快程度依賴于我們的algorithm的具體寫法。在本文中我們從理論山解釋了parallel comuting的limitations纽甘。同時(shí)展示了怎么借用comsol的batch sweep來improving performance when you reach these limits.

Amdahl’s and Gustafson-Barsis’ laws

算法分為serial algorithm和parallel algorithm亥鬓。通過增加計(jì)算單元（也叫作process或者threads），可以加快paralle algorithm的速度，但是對(duì)于serial algorithm 無效拢肆。我們實(shí)際中寫的algorithm大約是兩種algorithm的一種混合。假定代碼中parallel code 占比為 $\varphi$ ,則serial algorithm為（ $1-\varphi$ ）靖诗」郑考慮計(jì)算時(shí)間 $T(P)$ ，P代表進(jìn)程（Process）數(shù)目呻畸。當(dāng) $P=1$ 時(shí)移盆，計(jì)算時(shí)間記做 $T(1)$ ,那么當(dāng)active process 為P時(shí)，計(jì)算時(shí)間為： $T(P)=T(1) \cdot(\varphi / P+(1-\varphi))$ 伤为，那么相應(yīng)的speedup為： $S(P):=T(1) / T(P)=1 /(\varphi / P+(1-\varphi))$

Amdahl’s Law

For 100% parallelized code, the sky is the limit. 當(dāng) $\varphi<1$ 咒循，speedup會(huì)有一個(gè)limit：
$S_{\max }(\varphi):=\lim _{P \rightarrow \infty} S(P)=1 /(1-\varphi)$ 。比如如下圖所示： $S_{\max }(0.5)=2$

The speedup for increasing the number of processes for different fractions of parallelizable code

Gustafson-Barsis’ Law

Amdahl’s law assumes that the size of the problem is fixed. Yet, by assuming that the size of the problem increases with the number of added processes, then you are utilizing all the processes to an assumed level, and the speedup of the performed computations remains unbounded.

When taking into account that the size of the job normally increases with the number of available processes, our predictions are more optimistic.

The Cost of Communication

Gustafson-Barsis’ law implies that we are only restricted in the size of the problem we can compute绞愚，but sometimes communication is expensive. Let’s consider an overhead that is dominated by the communication and synchronization required in parallel processing, and model this as time added to the computation time.

In the case of no overhead, the result is as predicted by Amdahl’s law（last picture）, but when we start adding overhead, we see that something is happening.

Speedup with added overhead. The constant, c, is chosen to be 0.005. Different line represents the different function that describes the communication

For a quadratic function, the result is worse and, as you might recall from our earlier blog post on distributed memory computing, the increase of communication is quadratic in the case of all-to-all communication. Due to this phenomenon, we cannot expect to have a speedup on a cluster for, say, a small time-dependent problem when adding more and more processes. The amount of communication would increase faster than any gain from added processes. 不過我們此時(shí)考慮的是fixed size的problem叙甸，事實(shí)上，當(dāng)我們?cè)龃髉roblem的size的時(shí)候位衩， “slowdown” effect introduced through communication would be less relevant裆蒸。

Batch Sweeps in COMSOL Multiphysics

As our example model, we will use the electrodeless lamp, which is available in the Model Gallery. This model is small, at around 80,000 degrees of freedom, but needs about 130 time steps in its solution. To make this transient model parametric as well, we will compute the model for several values of the lamp power, namely 50 W, 60 W, 70 W, and 80 W.
On my workstation, a Fujitsu? CELSIUS? equipped with an Intel? Xeon? E5-2643 quad core processor and 16 GB of RAM, the following compute times are received:

Number of Cores	Compute Time per Parameter	Compute Time for Sweep
1	30 mins	120 mins
2	21 mins	82 mins
3	17 mins	68 mins
4	18 mins	72 mins

從上表可以看出，只是增加電腦利用的核數(shù)并不能增加速度糖驴，反而當(dāng)有3核改為4核之后速度變慢了僚祷。

We will now use the batch sweep functionality to parallelize this problem in another way: we will switch from data parallelism to task parallelism. We will create a batch job for each parameter value and see what this does to our computation times.

Simulations per day for the electrodeless lamp model. “4×1” means four batch jobs run simultaneously, using one core each.

從上圖可以看出佛致，當(dāng)我們把工作分成同時(shí)工作的四份，每份工作占用一個(gè)核辙谜，速度可以大大加快俺榆。

在我在自己的電腦上測(cè)試squareloop的工作的時(shí)候發(fā)現(xiàn)建立batch sweep確實(shí)也可以加快速度，我的電腦是4core装哆，16G of RAM. $1 \times 1$ 所用時(shí)間是3min41s, $2 \times 2$ 所用時(shí)間是2min20s罐脊， $4 \times1$ 所用時(shí)間是2min6s。加速效果并不是很明顯

Conclusion

在comsol中設(shè)置并行計(jì)算是個(gè)很復(fù)雜的問題蜕琴，就像怎么選擇求解器一樣萍桌。和要解決的問題，以及計(jì)算機(jī)的性能特點(diǎn)都很有關(guān)系凌简。

Selecting the right parallel configuration is not always easy, and it can be hard to know beforehand how you should “hybridize” your parallel computations. But as in many other cases, experience comes from playing around and testing, and with COMSOL Multiphysics, you have the possibility to do that. Try it yourself with different configurations and different models, and you will soon know how to set the software up in order to get the best performance out of your hardware.

?著作權(quán)歸作者所有,轉(zhuǎn)載或內(nèi)容合作請(qǐng)聯(lián)系作者

人面猴
序言：七十年代末上炎，一起剝皮案震驚了整個(gè)濱河市，隨后出現(xiàn)的幾起案子号醉，更是在濱河造成了極大的恐慌反症，老刑警劉巖，帶你破解...
沈念sama閱讀 217,907評(píng)論 6贊 506
死咒
序言：濱河連續(xù)發(fā)生了三起死亡事件畔派，死亡現(xiàn)場(chǎng)離奇詭異铅碍，居然都是意外死亡，警方通過查閱死者的電腦和手機(jī)线椰，發(fā)現(xiàn)死者居然都...
沈念sama閱讀 92,987評(píng)論 3贊 395
救了他兩次的神仙讓他今天三更去死
文/潘曉璐我一進(jìn)店門胞谈，熙熙樓的掌柜王于貴愁眉苦臉地迎上來，“玉大人憨愉，你說我怎么就攤上這事烦绳。” “怎么了配紫？”我有些...
開封第一講書人閱讀 164,298評(píng)論 0贊 354
道士緝兇錄：失蹤的賣姜人
文/不壞的土叔我叫張陵径密，是天一觀的道長(zhǎng)。經(jīng)常有香客問我躺孝，道長(zhǎng)享扔，這世上最難降的妖魔是什么？我笑而不...
開封第一講書人閱讀 58,586評(píng)論 1贊 293
?港島之戀（遺憾婚禮）
正文為了忘掉前任植袍，我火速辦了婚禮惧眠，結(jié)果婚禮上，老公的妹妹穿的比我還像新娘于个。我一直安慰自己氛魁，他們只是感情好，可當(dāng)我...
茶點(diǎn)故事閱讀 67,633評(píng)論 6贊 392
惡毒庶女頂嫁案：這布局不是一般人想出來的
文/花漫我一把揭開白布。她就那樣靜靜地躺著秀存，像睡著了一般捶码。火紅的嫁衣襯著肌膚如雪。梳的紋絲不亂的頭發(fā)上或链，一...
開封第一講書人閱讀 51,488評(píng)論 1贊 302
城市分裂傳說
那天宙项，我揣著相機(jī)與錄音，去河邊找鬼株扛。笑死，一個(gè)胖子當(dāng)著我的面吹牛汇荐，可吹牛的內(nèi)容都是我干的洞就。我是一名探鬼主播，決...
沈念sama閱讀 40,275評(píng)論 3贊 418
雙鴛鴦連環(huán)套：你想象不到人心有多黑
文/蒼蘭香墨我猛地睜開眼掀淘，長(zhǎng)吁一口氣：“原來是場(chǎng)噩夢(mèng)啊……” “哼旬蟋！你這毒婦竟也來了？” 一聲冷哼從身側(cè)響起革娄，我...
開封第一講書人閱讀 39,176評(píng)論 0贊 276
萬榮殺人案實(shí)錄
序言：老撾萬榮一對(duì)情侶失蹤倾贰，失蹤者是張志新（化名）和其女友劉穎，沒想到半個(gè)月后拦惋，有當(dāng)?shù)厝嗽跇淞掷锇l(fā)現(xiàn)了一具尸體匆浙，經(jīng)...
沈念sama閱讀 45,619評(píng)論 1贊 314
?護(hù)林員之死
正文獨(dú)居荒郊野嶺守林人離奇死亡，尸身上長(zhǎng)有42處帶血的膿包…… 初始之章·張勛以下內(nèi)容為張勛視角年9月15日...
茶點(diǎn)故事閱讀 37,819評(píng)論 3贊 336
?白月光啟示錄
正文我和宋清朗相戀三年厕妖，在試婚紗的時(shí)候發(fā)現(xiàn)自己被綠了首尼。大學(xué)時(shí)的朋友給我發(fā)了我未婚夫和他白月光在一起吃飯的照片。...
茶點(diǎn)故事閱讀 39,932評(píng)論 1贊 348
活死人
序言：一個(gè)原本活蹦亂跳的男人離奇死亡言秸，死狀恐怖软能，靈堂內(nèi)的尸體忽然破棺而出，到底是詐尸還是另有隱情举畸，我是刑警寧澤查排，帶...
沈念sama閱讀 35,655評(píng)論 5贊 346
?日本核電站爆炸內(nèi)幕
正文年R本政府宣布，位于F島的核電站抄沮，受9級(jí)特大地震影響跋核，放射性物質(zhì)發(fā)生泄漏。R本人自食惡果不足惜合是，卻給世界環(huán)境...
茶點(diǎn)故事閱讀 41,265評(píng)論 3贊 329
男人毒藥：我在死后第九天來索命
文/蒙蒙一了罪、第九天我趴在偏房一處隱蔽的房頂上張望。院中可真熱鬧聪全，春花似錦泊藕、人聲如沸。這莊子的主人今日做“春日...
開封第一講書人閱讀 31,871評(píng)論 0贊 22
一樁弒父案娃圆，背后竟有這般陰謀
文/蒼蘭香墨我抬頭看了看天上的太陽玫锋。三九已至，卻和暖如春讼呢，著一層夾襖步出監(jiān)牢的瞬間撩鹿，已是汗流浹背。一陣腳步聲響...
開封第一講書人閱讀 32,994評(píng)論 1贊 269
情欲美人皮
我被黑心中介騙來泰國(guó)打工悦屏，沒想到剛下飛機(jī)就差點(diǎn)兒被人妖公主榨干…… 1. 我叫王不留节沦，地道東北人。一個(gè)月前我還...
沈念sama閱讀 48,095評(píng)論 3贊 370
代替公主和親
正文我出身青樓础爬，卻偏偏與公主長(zhǎng)得像甫贯，于是被迫代替她去往敵國(guó)和親。傳聞我的和親對(duì)象是個(gè)殘疾皇子看蚜，可洞房花燭夜當(dāng)晚...
茶點(diǎn)故事閱讀 44,884評(píng)論 2贊 354